Architecture Design Diagrams

Last updated:

One Word, Four Bounded Contexts

Trust boundary

Customer means something different inside each context's model.

The word customer in four bounded contexts Four bounded contexts, each with its own model of a customer. To sales, a customer is a lead with a pipeline stage. To fulfillment, it is a delivery address. To billing, it is a payment history and a credit limit. To support, it is a ticket history. SALES Customer a lead with a pipeline stage FULFILLMENT Customer a delivery address BILLING Customer a payment history and a credit limit SUPPORT Customer a ticket history

Context Map

C4 · Context

Fulfillment's upstream contexts, and how it integrates with each.

Context map around a fulfillment context Three upstream contexts feed the fulfillment context, which is downstream of each. The sales context is an open host service, and fulfillment consumes its product and order API. The legacy warehouse is a big ball of mud, and fulfillment protects its model with an anticorruption layer toward it. The payment provider is external, and fulfillment is a conformist toward it, adopting its model as it is. Sales context Open Host Service Fulfillment context ACL toward warehouse Conformist toward payment provider Legacy warehouse Big Ball of Mud Payment provider external U D product and order API U D U D U = upstream, D = downstream. Arrows point from upstream to downstream.

Order Aggregate

Structure

Changes go through the root, and other aggregates are referenced by id.

An order aggregate and its reference to a customer aggregate The order aggregate contains the order, which is the root, and its order lines. The root enforces three invariants: the total equals the sum of the lines, no changes once confirmed, and no confirming an empty order. Outside the boundary, the order references the separate customer aggregate by id only, through a CustomerId. ORDER AGGREGATE Order aggregate root Root enforces: total = sum of lines no changes once confirmed no confirming an empty order OrderLine OrderLine OrderLine CustomerId by id only Customer aggregate separate

Offset and Cursor Pages Under an Insert

Chart

An insert shifts offset pages, while a cursor keeps its place.

An inserted item shifting offset pagination but not cursor pagination Page 1, fetched with offset 0 and limit 3, returns items A, B, and C. Before page 2 is fetched, a new item X is inserted at the front of the collection, so it now reads X, A, B, C, D, E, F. Page 2 with offset 3 returns C, D, and E, so C is repeated. Page 2 with a cursor positioned after C returns D, E, and F, with nothing repeated or skipped. Page 1 offset 0, limit 3 A B C Before page 2, X is inserted first X A B C D E F Page 2, offset offset 3, limit 3 C D E repeated Page 2, cursor after C, limit 3 D E F offset 3

N+1 Queries and Batching

Flow

Per-order customer lookups, against one batched load through a DataLoader.

N+1 queries from naive resolvers compared with DataLoader batching Two panels for a query asking for 50 orders with their customers. With naive resolvers, the GraphQL server runs one query for the orders and then one query per order for its customer, 51 queries in total. With a DataLoader, the server runs one query for the orders and one batched query for every customer id collected while resolving that level, 2 queries in total. NAIVE RESOLVERS GraphQL server Database orders: 1 query customer: 1 query per order 51 queries DATALOADER GraphQL server Database orders: 1 query customers: 1 batched query for all ids 2 queries

gRPC Communication Patterns

C4 · Dynamic

Unary, server streaming, client streaming, and bidirectional calls.

The four gRPC communication patterns Four panels, each with a client and a server. Unary: one request and one response. Server streaming: one request and a stream of responses. Client streaming: a stream of requests and one response. Bidirectional streaming: both sides send streams of messages independently over one call. UNARY Client Server One request, one response SERVER STREAMING Client Server One request, a stream of responses CLIENT STREAMING Client Server A stream of requests, one response BIDIRECTIONAL Client Server Both sides stream independently request message response message time runs downward

Balancing gRPC Connections and Calls

Flow

Connection-level balancing pins each client to one pod.

Layer 4 connection balancing compared with call-level balancing for gRPC Two panels. With layer 4, connection-level balancing, client A holds one connection that the load balancer sends to pod 1, and client B holds one connection sent to pod 2, so each pod receives all of one client's calls and pod 3 stays idle, including after scale-out. With layer 7 or client-side, call-level balancing, client A's single connection reaches a proxy or balancing client that spreads its calls across pods 1, 2, and 3. LAYER 4: CONNECTION-LEVEL Client A Client B LB Pod 1 Pod 2 Pod 3 all of A's calls all of B's calls idle, including after scale-out LAYER 7 OR CLIENT-SIDE: CALL-LEVEL Client A proxy or client Pod 1 Pod 2 Pod 3 one connection

gRPC from a Browser

Flow

gRPC-Web or JSON transcoding translates to native gRPC before the service.

Two ways for a browser to reach a gRPC service Two paths. In the first, the browser calls with gRPC-Web over HTTP/1.1 or HTTP/2 to gRPC-Web support, provided by Envoy or by the server itself, which calls the service with gRPC. In the second, the browser calls with REST and JSON to JSON transcoding, running in-process or in a proxy such as grpc-gateway, which calls the service with gRPC. Browser gRPC-Web support Envoy, or the server itself Service gRPC-Web HTTP/1.1 or HTTP/2 gRPC Browser JSON transcoding in-process, or a proxy such as grpc-gateway Service REST + JSON gRPC

The OpenAPI Spec Pipeline

Flow

Build-time documents per audience, checked before the public one ships.

How OpenAPI documents flow from application code through build-time generation, linting, and a breaking-change diff to a spec archive and a developer portal Application code assigns each endpoint to a group. The build generates one OpenAPI document per group: a public document with the customer surface and an internal document with the full surface. The public document is linted and then diffed against the last published version. A breaking change fails the build unless someone acknowledges it. A document that passes is stored in the spec archive, and the developer portal, a static site behind a CDN, renders documentation from the archived specs. The internal document goes only to an internal explorer in development and staging. Application code Each endpoint assigned to a group Build One document per group Public document Customer surface Internal document Full surface Lint Rule checks Diff Against last published Build fails Unless break is acknowledged break Spec archive One spec per release passes Developer portal Static site, CDN Internal explorer Development and staging only document moves on breaking change stops the pipeline

The Lost Update

C4 · Dynamic

Client A's full-state write silently overwrites client B's change.

Two clients losing an update through full-state writes Client A gets the profile and receives email a@x and phone 1. Client B gets the same profile. B changes the phone and puts email a@x, phone 2, and the server stores phone 2. A then changes the email and puts email b@x with the stale phone 1, and the server stores email b@x, phone 1. B's phone change is gone. Client A Server Client B GET /profile {email: a@x, phone: 1} GET /profile {email: a@x, phone: 1} PUT {email: a@x, phone: 2} B changes phone stored: phone 2 PUT {email: b@x, phone: 1} A changes email, resending stale phone stored: email b@x, phone 1 B's phone change is gone

Server-Side Apply Field Ownership

Structure

Each field has a manager, and changing another manager's field is a conflict.

Field ownership on a Kubernetes Deployment under server-side apply A Deployment object with two fields. A human applying through kubectl owns the web container's image field. An autoscaler controller owns the replicas field. Each manager sets its own field independently. When kubectl applies a different replicas value, the API server rejects the whole apply with 409 Conflict because the autoscaler owns that field. The caller either drops the field or resubmits with force-conflicts to take ownership. kubectl human field manager autoscaler controller field manager DEPLOYMENT containers[web].image managed by kubectl replicas managed by autoscaler sets replicas 409 Conflict whole apply rejected Drop the field, or resubmit with --force-conflicts.

Tenancy Models

C4 · Deployment

Silo, pool, and the two bridge shapes, by what tenants share.

Silo, pool, and bridge tenancy models for three tenants Four panels for tenants A, B, and C. Silo: each tenant has its own application and its own database. Pool: one application and one database serve all three. Bridge along layers: one shared application tier writes to a separate database per tenant. Bridge along tenants: tenants A and B share a pooled application and database, and tenant C has a dedicated application and database. SILO Nothing shared below onboarding, identity, and operations App A DB A App B DB B App C DB C POOL All compute, storage, and messaging shared App: A, B, C DB: A, B, C BRIDGE: ALONG LAYERS Shared app tier, one database per tenant App: A, B, C DB A DB B DB C BRIDGE: ALONG TENANTS Most tenants pooled, a few in dedicated silos App: A, B DB: A, B App C DB C

Deployment Stamps

C4 · Deployment

A router reads the tenant catalog and sends each request to its stamp.

Tenants routed across deployment stamps through a tenant catalog A request carrying a tenant identity reaches a global router or gateway, which looks up the tenant in a tenant catalog mapping each tenant to a stamp. The router forwards the request to one of three stamps. Stamp 1 is a pool with shared app and database for tenants A, B, and C. Stamp 2 is a pool for tenants D and E. Stamp 3 is a silo serving tenant F only, with the same code on dedicated infrastructure. request carrying a tenant identity Global router or gateway Tenant catalog tenant → stamp Stamp 1 pool Tenants A, B, C shared app, DB Stamp 2 pool Tenants D, E shared app, DB Stamp 3 silo Tenant F only same code, dedicated infra

Tenant Context Across Hops

Flow

Where the tenant ID lives at each hop, from token to database.

Carrying tenant context from an HTTP request through a queue to the database An HTTP request carries the tenant ID as a token claim. The API holds it in a request-scoped TenantContext. It publishes to a queue with the tenant ID in a message header. The worker sets its own TenantContext from that header before the handler runs, and the database receives it as the transaction-local app.tenant_id setting. The API's outbound call to a downstream service carries the tenant in a header or token, and the downstream service resolves it again. HTTP request token claim tenant_id API TenantContext request scope Queue message header tenant_id Worker TenantContext set from header before the handler Database transaction-local app.tenant_id Outbound call header or token Downstream service resolves the tenant again

Agreed Risks on the Diagram

C4 · Container

Availability risks pinned to the parts they apply to.

Availability risks placed on an architecture diagram after risk storming An API gateway calls an order service and a payment service, and the order service also calls the payment service. The order service writes to an orders database that is a single instance, and the payment service uses a payment provider that is one vendor. Two agreed availability risks are pinned to the diagram: a 6 on the orders database, because it has no replica or failover, and a 9 on the payment provider, because there is no fallback provider and an outage stops checkout. API gateway Order service Payment service Orders database single instance Payment provider one vendor 6 no replica or failover 9 no fallback provider, outage stops checkout

Test Shapes

Chart

Pyramid, ice cream cone, trophy, and honeycomb, by where tests concentrate.

Four test suite shapes Four shapes, each a stack of test scopes from bottom to top, with width showing relative emphasis. The pyramid is widest at unit tests, narrower at service tests, and narrowest at UI or end-to-end tests. The ice cream cone inverts it, with few unit tests and mostly manual and end-to-end tests. The trophy has static analysis at the base, then unit tests, then its widest band of integration tests, and few end-to-end tests. The honeycomb has few tests of implementation detail, mostly integration tests of each service, and fewest integrated tests across services. PYRAMID Unit Service UI / E2E ICE CREAM CONE Unit Service Manual / E2E TROPHY Static Unit Integration E2E HONEYCOMB Impl. detail Integration Integrated Width shows relative emphasis, not a test count. The ice cream cone is the antipattern the others avoid.

Consumer-Driven Contract Flow

Flow

The consumer publishes a pact, and the provider verifies it through a broker.

Consumer-driven contract testing with Pact and a Pact Broker In the consumer build, consumer tests run against a Pact mock provider and produce a pact file recording the requests made and the response fields relied on. The consumer publishes it to the Pact Broker. In the provider build, the provider starts on a real port, Pact replays each recorded request from the broker and checks the responses match, and the provider publishes the verification result back to the broker. Before either side releases a version, a can-i-deploy check against the broker confirms compatibility. CONSUMER BUILD Consumer tests run against a Pact mock provider Pact file requests made, response fields relied on Pact Broker publish PROVIDER BUILD Provider on a real port Pact replays requests each recorded request, checks the responses match pact verification result can-i-deploy before either side releases

Tail Latency Under Fan-Out

Flow

Waiting on 100 servers makes a 1% tail the common case.

A user request fanning out to 100 servers with a one-second p99 A user request reaches a front end, which calls 100 servers in parallel, each with a p99 latency of one second, and waits for all of them. The probability that all 100 beat one second is 0.99 to the power 100, about 0.37, so about 63 percent of user requests take longer than one second. User request Front end Server 1 p99 = 1 s Server 2 p99 = 1 s ⋮ Server 99 p99 = 1 s Server 100 p99 = 1 s The front end waits for all 100. P(all beat 1 s) = 0.99¹⁰⁰ ≈ 0.37 About 63% of user requests take longer than 1 s.

Latency Against Utilization

Chart

Response time climbs steeply as a resource nears full use.

Average response time as a multiple of service time, by utilization A curve from the simplest single-server queueing model, where average response time equals service time divided by one minus utilization. The curve is nearly flat at low utilization, reaches twice the service time at 50 percent, five times at 80 percent, and ten times at 90 percent, and climbs steeply beyond that toward full use. 0% 25% 50% 75% 100% 0× 5× 10× 15× 20× 50%: 2× service time 80%: 5× service time 90%: 10× service time utilization average response time, as a multiple of service time

Why Scaling Out Isn't Linear

Chart

Contention caps throughput, and coherency costs can make it fall.

Throughput against instance count under the Universal Scalability Law Three illustrative curves of throughput as instances are added. Linear scaling rises in proportion to instance count. With contention for a shared resource, throughput flattens toward a ceiling no matter how many instances are added. With coherency costs as well, throughput peaks and then falls as more instances are added. The curves are illustrative, not measured. instances throughput Linear what adding instances promises Contention shared resources set a ceiling Plus coherency staying consistent costs more with every instance Illustrative curves, not measurements.

Routing Changes to a Review Level

Flow

Two questions send a change to self-service, peer, or board review.

Routing a proposed architecture change to a level of review A proposed architecture change meets a first question: does it use approved patterns and technology, with consequences that stay within one team? If yes, it goes to self-service, with automated checks and the team recording the decision. If no, a second question asks whether it crosses team or system boundaries, introduces new technology, or affects security, compliance, or shared data. If no, it goes to peer review by another team or a community of practice. If yes, it goes to board-level review, by a review board or through broad advice from affected teams and experts. Proposed architecture change Uses approved patterns and technology, and its consequences stay within one team? yes Self-service Automated checks, team records the decision no Crosses team or system boundaries, introduces new technology, or affects security, compliance, or shared data? no Peer review Another team or community of practice reviews the design yes Board-level review Review board, or broad advice from affected teams and experts

TOGAF Architecture Development Method

Flow

Eight phases in a cycle, with Requirements Management at the center.

The TOGAF Architecture Development Method cycle A Preliminary phase leads into a cycle of eight phases: A, Architecture Vision; B, Business Architecture; C, Information Systems Architectures; D, Technology Architecture; E, Opportunities and Solutions; F, Migration Planning; G, Implementation Governance; and H, Architecture Change Management, which leads back to A. Requirements Management sits at the center and exchanges requirements with every phase. A. Architecture Vision B. Business Architecture C. Information Systems Architectures D. Technology Architecture E. Opportunities and Solutions F. Migration Planning G. Implementation Governance H. Architecture Change Management Requirements Management Preliminary

TCO Sensitivity to Staffing

Chart

The staffing assumption flips which option costs less over three years.

Three-year cost of self-managed and managed streaming against operations staffing An illustrative chart from the worked example. The managed service's three-year total is flat at $675K. The self-managed total depends on the operations effort it needs: $910K with one full engineer, and $640K with half an engineer. The self-managed line crosses the managed line between those two staffing levels, so the staffing assumption decides which option is cheaper. 0.5 0.75 1.0 $600K $700K $800K $900K $1000K 1.0 engineer: $910K 0.5 engineer: $640K self-managed operations effort, in engineers three-year total Self-managed Managed service $675K Illustrative figures from the example.

NPV Sensitivity and Break-Even

Chart

NPV against annual benefit, crossing zero at the break-even value.

Net present value against annual benefit for the example investment A straight line of NPV at 10 percent over five years against the annual benefit of a $100K investment. At $25K a year NPV is minus $5.2K, at $40K it is $51.6K, and at $55K it is $108.5K. The line crosses zero at the break-even benefit of about $26.4K a year. $20K $30K $40K $50K $60K −$40K $0K $40K $80K $120K $25K: −$5.2K $40K: $51.6K $55K: $108.5K break-even ≈ $26.4K annual benefit NPV at 10% over 5 years, for a $100K investment NPV below zero

Found this useful? Share it:

Share on LinkedIn