Infrastructure Diagrams

Last updated:

The Plan/Apply Loop

Flow

Code, state record, and live resources compared into a reviewed plan.

How an IaC tool turns definition files into changes The tool compares the definition files, which describe what should exist, with its state record, which maps each definition to a real resource ID. Terraform, OpenTofu, and Bicep also read the live infrastructure's current configuration on every plan; Pulumi and CloudFormation read it only on request. Bicep keeps no state record. The comparison produces a plan of create, update, replace, and delete actions. After review, apply changes the live resources and updates the state record. State record definition → real resource ID (Bicep needs none) Definition files what should exist Live infrastructure read through the cloud API Compare desired vs. current Plan create · update · replace · delete Apply review changes resources updates the record every tool (Bicep has no record) every plan for Terraform, OpenTofu, and Bicep; on request for Pulumi and CloudFormation

Mutable vs. Immutable Servers

Flow

Patching running servers in place vs. replacing them from a new image.

How a change reaches servers under mutable and immutable infrastructure On the mutable side, a configuration tool pushes version 2 onto three running servers. Two upgrade cleanly and one is only partly upgraded, and each server keeps every change ever applied to it. On the immutable side, an image build produces a version 2 image, three new servers start from it, and then the three old version 1 servers are removed. Mutable: change the servers you have Configuration tool pushes v2 to each server Server 1 v1 → v2 Server 2 v1 → v2 Server 3 partly v2 each server keeps every change ever applied to it Immutable: replace them Image build produces image v2 Server 4 started from v2 Server 5 started from v2 Server 6 started from v2 the v1 servers they replace Old server (v1) Old server (v1) Old server (v1) removed after the new servers start

Two Runs, One State

C4 · Dynamic

Without a lock, the last write drops the other run's resources.

Two runs applying against one state, without and with a lock Without a lock, run A and run B both read state version 5. Run A creates web-1 and writes version 6 listing it. Run B creates db-1 and writes its own version 6 listing only db-1, overwriting run A's record, so web-1 keeps running but is no longer tracked. With a lock, run A takes the lock and reads version 5. Run B's attempt to take the lock is refused. Run A writes version 6 and releases the lock, then run B takes it, reads version 6, and writes version 7, which lists both web-1 and db-1. Without a lock Run A State Run B reads v5 reads v5 writes v6: web-1 writes v6: db-1 Record lists db-1 only web-1 still runs, but the tool no longer tracks it With a lock Run A State Run B locks, reads v5 lock refused writes v6, unlocks locks, reads v6 writes v7, unlocks Record lists web-1 and db-1 Run B started from run A's record read or write of the state record a write that loses data, or a refused lock

Breaking a Dependency Cycle

Dependency graph

Two security groups that reference each other, before and after separating the rule.

A dependency cycle between two security groups, and how a separate rule resource breaks it With rules written inline, the app security group depends on the db security group and the db security group depends on the app security group, so the dependency graph has a cycle and neither can be created first. With the rule as a separate resource, both security groups depend on nothing, and the rule depends on both, so the tool creates the two groups first and the rule last. Rules inline: a cycle app group rule names db db group rule names app neither can be created first Terraform stops with a cycle error Rule separate: no cycle app group no rules db group no rules ingress rule db allows app Created in order: both groups, then the rule depends on (must be created after) a dependency that closes a cycle

Push vs. Pull Delivery

Flow

A pipeline pushes each merged change, while an agent pulls and reconciles.

Push-based pipeline delivery compared with pull-based GitOps delivery In push delivery, an event in the Git repository, such as a merge or a pull-request comment, triggers a CI pipeline, which holds the cloud credentials and applies the change to the resources. Between runs nothing reconciles, and drift shows up only if a scheduled check is set up. In pull delivery, an agent pulls the desired state from the Git repository on its own schedule, applies it to the resources directly or through an operator, and keeps observing their actual state, comparing the two on every interval. Push: a pipeline applies Git repository CI pipeline runs when an event fires Cloud resources an event triggers a run applies the change Between runs, nothing reconciles Pull: an agent reconciles Git repository GitOps agent runs on a loop Cloud resources pulls desired state applies, or via an operator observes actual state Compares desired and actual state on every interval the component that applies changes

Three Developer Environment Models

Structure

Which layers each model shares and which each developer gets a copy of.

Shared and per-developer layers in three developer environment models In the shared data model, each developer has their own application stack, and the data stores and the network are shared by all developers. In the dedicated model, each developer has their own application stack, data stores, and network. In the shared network, personal data stores model, each developer has their own application stack and data stores, and the network is shared. Application Data stores Network Shared data dev 1 dev 2 dev 3 shared shared Dedicated dev 1 dev 2 dev 3 dev 1 dev 2 dev 3 dev 1 dev 2 dev 3 Shared network, personal data dev 1 dev 2 dev 3 dev 1 dev 2 dev 3 shared one copy used by every developer a separate copy per developer

Where Governance Controls Sit

Flow

Two paths a change can take, and which controls each path passes.

Where preventive, detective, and corrective controls sit on the paths a change can take A change made through infrastructure as code passes pipeline checks, then a deployment-service hook if the tool deploys through one, then the cloud API, where organization policy applies, before it reaches the running resources. A change made by hand in the console, a CLI, or a script bypasses the pipeline checks and the hook and goes straight to the cloud API, where organization policy still applies to every account it covers. Pipeline checks, hooks, and organization policy are preventive. Detective rules evaluate the running resources they cover after the fact, and corrective actions triggered by their findings fix or revert the resources. Change through IaC Pipeline checks sees this pipeline Service hook if the tool uses one Cloud API organization policy sees covered accounts Running resources Change by hand console, CLI, script skips the pipeline checks and the hook Detective rules sees covered resources Corrective action fixes or reverts evaluates fixes preventive: stops a change before it takes effect detective and corrective: act after the resource exists

Traffic Shift by Strategy

Chart

Share of traffic on the new version over a release, for three strategies.

Share of traffic on the new version over time for rolling, blue-green, and canary deployments Three small charts, each plotting the share of traffic on the new version from 0 to 100 percent over the course of one release. In a rolling deployment the share rises in even steps as each batch of instances is replaced. In a blue-green deployment the share stays at zero while the idle green environment is tested, then jumps to 100 percent at the switch. In a canary release the share rises in uneven steps of 5, 25, 50, and 100 percent, with a pause at each step while the new version is watched. Rolling 100% 50% 0% one batch at a time time Blue-green 100% 50% 0% green tested, no users switch time Canary 100% 50% 0% 5% 25%, then 50% time Share of traffic on the new version during one release. Step sizes vary by tool.

RTO and RPO on a Timeline

Chart

RPO reaches back from the disaster, RTO forward from it.

The recovery point objective and recovery time objective on a timeline around a disaster A timeline with three marked moments: the last recovery point, the disaster, and service restored. The span from the last recovery point to the disaster is the data that is lost, which the recovery point objective limits. The span from the disaster to service restored is the time the service is unavailable, which the recovery time objective limits. RPO limits this data written since the last recovery point is lost RTO limits this the service is unavailable time Last recovery point backup or replicated write Disaster service interrupted Service restored in the recovery site

Found this useful? Share it:

Share on LinkedIn