Observability Diagrams

Last updated:

A Checkout Trace as a Waterfall

Chart

One slow checkout request's spans on a shared timeline.

Waterfall view of one checkout trace across three services Seven spans on a timeline from 0 to 3.4 seconds, indented under their parents. POST /checkout in the checkout service takes 3,400 ms. Under it, SELECT cart takes 40 ms, POST /orders in the orders service takes 180 ms with an INSERT order child of 120 ms, and POST /payments in the payments service takes 3,100 ms. Under payments, the first POST /charge client span to the external payment gateway runs for 2,000 ms and ends in Error with a timeout, and a retried POST /charge succeeds in 1,000 ms. SPAN SERVICE 0 1 s 2 s 3 s POST /checkout checkout 3,400 ms SELECT cart checkout, client 40 ms POST /orders orders 180 ms INSERT order orders 120 ms POST /payments payments 3,100 ms POST /charge (1st) payments, client 2,000 ms · Error: timeout POST /charge (retry) payments, client 1,000 ms · Ok checkout orders payments span status Error

Trace Context Propagation

Flow

The traceparent header carried from hop to hop, and dropped.

The traceparent header carried across services, and a hop that drops it Top: checkout sends POST /payments with a traceparent header carrying trace ID 4bf9…4736 and its client span ID 00f0…02b7. Payments starts a server span whose parent is 00f0…02b7 in the same trace, then calls the payment gateway with a new traceparent header carrying the same trace ID and its own client span ID b7ad…3331. Bottom: the same call passes through a proxy that strips unknown headers, so payments receives no traceparent and starts a new trace 9e1c…77a0, and one request is recorded as two unrelated traces. CONTEXT PROPAGATED checkout trace 4bf9…4736 client span 00f0…02b7 traceparent 00-4bf9…4736-00f0…02b7-01 payments trace 4bf9…4736 new server span, parent 00f0…02b7 client span b7ad…3331 traceparent 00-4bf9…4736-b7ad…3331-01 payment gateway external sends no spans CONTEXT DROPPED checkout trace 4bf9…4736 traceparent proxy strips unknown headers no traceparent payments starts trace 9e1c…77a0 One request, recorded as two unrelated traces.

A CPU Flame Graph

Chart

The checkout service's CPU samples, stacked by call path.

CPU flame graph for the checkout service Stacked boxes, each a function sitting on top of the function that called it, with width proportional to the share of CPU samples. At the bottom, all samples. Above it, garbage collection takes 8 percent and a thread pool worker 92 percent. The worker runs CheckoutController.Post at 88 percent. Post calls Cart.Validate at 12 percent, JsonSerializer.Serialize at 46 percent, and TaxCalculator.Compute at 22 percent. Above those, Regex.Match takes 9 percent under Cart.Validate, PropertyInfo.GetValue 16 percent and Utf8JsonWriter.WriteString 28 percent under JsonSerializer.Serialize, and TaxRateCache.Lookup 15 percent under TaxCalculator.Compute. all samples (100%) GC 8% ThreadPool worker 92% CheckoutController.Post 88% Cart.Validate JsonSerializer.Serialize 46% TaxCalculator.Compute 22% Regex.Match PropertyInfo.GetValue Utf8JsonWriter.WriteString TaxRateCache.Lookup width = share of CPU samples, sorted by name (not a timeline) caller below, callee above application code framework and runtime

A Multiwindow Burn-Rate Alert

Chart

A 1-hour and a 5-minute window watching one bad deploy.

Error rate over a 1-hour and a 5-minute window during a 30-minute incident A chart over 100 minutes. The error rate is 5 percent from minute 10 to minute 40, when a fix is deployed, and zero otherwise. The threshold for a burn rate of 14.4 on a 99.9 percent SLO is 1.44 percent errors. The 5-minute window's error rate rises to 5 percent by minute 15 and falls to zero by minute 45, staying above the threshold until about minute 44. The 1-hour window's error rate climbs to 2.5 percent by minute 40, holds until minute 70, and falls to zero at minute 100, crossing the threshold at about minute 27 and staying above it until about minute 83. Below the chart, an alert on the 1-hour window alone fires from minute 27 to minute 83, 43 minutes past the fix. An alert requiring both windows fires from minute 27 to minute 44, about 4 minutes past the fix. 0% 2% 4% 6% 0 20 40 60 80 100 min burn rate 14.4 threshold (1.44% errors) fix deployed actual error rate 5-minute window 1-hour window 1 h alone fires 43 min past the fix 1 h and 5 min fires 4 min past the fix error rate

Agent-to-Gateway Collectors

C4 · Deployment

Agents route spans by trace ID so each trace reaches one gateway.

Agent Collectors routing spans by trace ID to a pool of gateway Collectors Three hosts run the checkout, orders, and payments services, each with the OpenTelemetry SDK and an agent Collector running a load-balancing exporter. Spans of trace 4bf9…4736 come from all three hosts and are all routed to gateway Collector 1. Spans of trace 9e1c…77a0 also come from all three hosts and are all routed to gateway Collector 2. Each gateway runs tail sampling on complete traces and sends the traces it keeps to the tracing backend. HOST 1 checkout [OTel SDK] agent [load balancing] HOST 2 orders [OTel SDK] agent [load balancing] HOST 3 payments [OTel SDK] agent [load balancing] gateway Collector 1 tail sampling gateway Collector 2 tail sampling tracing backend stores kept traces spans of trace 4bf9…4736 spans of trace 9e1c…77a0 Agents hash each span's trace ID to pick a gateway, so every span of a trace reaches the same one.

Found this useful? Share it:

Share on LinkedIn