LLM Systems Diagrams

Last updated:

Chain-of-Thought, Self-Consistency, Tree of Thoughts

Structure

One path, several paths with a vote, and a search tree.

The shapes of three reasoning techniques Chain-of-thought: one linear path from the question through several steps to an answer, so an early mistake carries through. Self-consistency: several independent reasoning paths from the same question, each reaching an answer, followed by a vote on the answer most paths agree on, so a wrong path is outvoted if most avoid it. Tree of Thoughts: from the problem, several candidate next steps are proposed and evaluated; a weak branch is abandoned, a dead end is backtracked from, and the promising branch is explored further to a solution. Answers shown are illustrative. Chain-of-thought one linear path Self-consistency several paths, then a vote Tree of Thoughts search with evaluation Question Step 1 Step 2 Step 3 Answer An early mistake carries through. Question 42 42 38 42 40 Vote: 42 A wrong path is outvoted if most avoid it. Problem A B C abandoned A1 dead end: backtrack B1 Solution Each candidate step is evaluated. Weak branches are abandoned.

MCP Hosts, Clients, and Servers

C4 · Container

One client per server inside the host; only the host talks to the model.

MCP host, client, and server architecture A host application such as a chat app, IDE, or agent runtime owns the conversation with the model API and sends it the conversation, including tool results. Inside the host, one client connects to each server: client A to a local filesystem server over stdio, clients B and C to remote ticketing and internal API servers over Streamable HTTP. Each server reaches its own backend: local files, a ticketing API, or internal services. The model never talks to the servers directly. Model API [Model provider] never talks to servers Host application [Chat app, IDE, or agent runtime] owns the conversation, can gate and log every call conversation, including tool results Client A [one per server] Client B [one per server] Client C [one per server] stdio Filesystem server [Local child process] Local files Streamable HTTP Ticketing server [Remote service] Ticketing API Streamable HTTP Internal API server [Remote service] Internal services

MCP Authorization Flow

C4 · Dynamic

From an unauthenticated request to an audience-bound bearer token.

The MCP OAuth authorization sequence 1: the MCP client sends a request with no token. 2: the MCP server answers 401 with a WWW-Authenticate header naming its resource metadata URL and required scope. 3: the client fetches the protected resource metadata, which names the authorization server to use. 4: the client fetches the authorization server's metadata. 5: the client identifies itself with a client ID metadata document URL, a pre-registered ID, or dynamic client registration. 6: the client sends an authorization request with a PKCE challenge and the MCP server's URI as the resource, and the user signs in and consents in a browser. 7: the authorization server returns an authorization code with its issuer, which the client checks. 8: the client requests a token with the code, PKCE verifier, and resource. 9: the authorization server issues an access token whose audience is this MCP server. 10: the client sends requests to the MCP server with that bearer token. MCP client MCP server Authorization server 1 · Request, no token 2 · 401 + WWW-Authenticate resource_metadata URL, required scope 3 · Fetch protected resource metadata which authorization server to use 4 · Fetch authorization server metadata 5 · Identify the client metadata document URL, pre-registered ID, or dynamic registration 6 · Authorization request PKCE challenge + resource = the MCP server's URI user signs in and consents in a browser 7 · Authorization code + issuer, checked by the client 8 · Token request code + PKCE verifier + resource 9 · Access token, audience = this MCP server 10 · Requests with Authorization: Bearer token

Two RAG Pipelines Sharing One Index

Flow

Offline indexing and per-request querying, meeting at one index.

The indexing and query pipelines of a RAG system Indexing runs offline whenever content changes: source documents are parsed and cleaned, chunked, embedded, and written to an index holding vectors, text, metadata, and access rules. The query pipeline runs per request: the user question is rewritten and embedded, candidates are retrieved from the index with vector and keyword search filtered by metadata and user permissions, reranked to the top passages, assembled into a prompt with instructions and the question, and sent to the model, which answers with citations. INDEXING · OFFLINE, WHENEVER CONTENT CHANGES QUERY · ONLINE, PER REQUEST Source documents Parse and clean keep structure Chunk Embed Index [Shared store] vectors, text, metadata, access rules User question Rewrite and embed the query Retrieve candidates vector search + keyword search, filtered by metadata and user permissions searched Rerank Top passages Prompt instructions + passages + question Model Answer with citations

Bi-Encoder and Cross-Encoder

Structure

Encoding separately for search versus reading together to score.

Bi-encoder compared with cross-encoder Left: a bi-encoder encodes the document and the query separately; document vectors are computed once, and query and document vectors are compared by similarity, which makes search across millions of chunks fast. Right: a cross-encoder reads the query and one document together and outputs a relevance score; it is more accurate but produces no reusable embedding and must run once per query-document pair, too slow for a whole collection. Bi-encoder used for retrieval embeddings Cross-encoder used for reranking Document Encoder Doc vector computed once Query Encoder Query vector Similarity Query and documents are encoded separately, so document vectors are reused and search across millions of chunks is fast. Query + one document read together Cross-encoder Relevance score More accurate, but no reusable embedding. Runs once per query-document pair, too slow for a whole collection.

Reading Fine-Tuning Loss Curves

Chart

Four loss-curve shapes and the response to each.

Four fine-tuning loss curve patterns Four small charts of loss over training steps. Learning: training and validation loss both fall; continue. Overfitting: training loss keeps falling while validation loss turns and rises; stop earlier or add data. Flat from the start: both stay flat; check that labels cover the responses, then raise the learning rate. Spiking or diverging: training loss falls, then spikes and climbs; lower the rate or add warmup. Shapes are illustrative. Learning steps Continue Overfitting steps Stop earlier or add data Flat from the start steps Check masking, raise rate Spiking or diverging steps Lower rate, add warmup training loss validation loss Shapes are illustrative.

One Agent Cycle Across the Network

C4 · Dynamic

Tools run locally, inference runs remotely, and context crosses on every turn.

One agent loop cycle split across local execution and remote inference Steps 1, 2, 5, 6, and 9 run on the user's machine: the user provides a goal, the runtime assembles the initial context from the system prompt and goal, executes a tool locally such as reading file X from disk, appends the tool result to the context, and executes the next tool. Steps 3, 4, 7, and 8 run on the provider's servers: the model reasons about the goal, returns a tool call, sees the file contents and reasons about the next step, and returns the next tool call or a final response. Each crossing is an HTTPS request; the red rightward crossings are data leaving the machine, carrying the whole conversation so far. LOCAL EXECUTION · USER'S MACHINE REMOTE INFERENCE · PROVIDER HTTPS 1 · User provides goal 2 · Assemble initial context system prompt + user goal 3 · Model reasons about goal 4 · Model returns a tool call “read file X” 5 · Execute tool locally read file X from disk 6 · Append tool result to the context 7 · Model sees file contents, reasons about next step 8 · Returns next tool call or a final response 9 · Execute next tool locally the cycle repeats Red: a request leaving the machine, carrying the whole conversation so far. Blue: a response coming back.

One Context Window, No Privilege Levels

Trust boundary

Untrusted sources share one context; its output reaches downstream sinks.

Instructions and data share one context window User messages, retrieved chunks, tool results, fetched web pages, file contents, and output from another agent are untrusted sources that all flow into one context window. The system prompt enters the same window; it is trusted, but the architecture gives it no privilege over the rest. The context window has no privilege levels. The model output derived from it is therefore untrusted, and it reaches downstream sinks: rendering in a browser (XSS), a shell or eval (remote code execution), a SQL query (injection), a tool call with real-world effect, and a URL the client fetches (data exfiltration). UNTRUSTED SOURCES User message Retrieved chunks Tool results Fetched web page File contents Another agent System prompt trusted, but not privileged by the architecture Context window [One token sequence] system prompt, user input, documents, and tool results side by side, with no privilege levels Model output [Untrusted] derived from every input above DOWNSTREAM SINKS Browser render XSS Shell or eval remote code execution SQL query injection Tool call real-world effect Client-fetched URL data exfiltration Dashed red: untrusted. Every source can try to steer the model, and every sink receives text an attacker may have shaped.

Batch Cost: One Session Versus Isolated Tasks

Chart

Cumulative input cost across a 30-item batch, three ways.

Cumulative input cost of a batch run three ways Cumulative input tokens billed across a 30-item batch. One long session grows quadratically, because each request resends every earlier output. The same session with prompt caching still grows quadratically, but the reused history is billed at a tenth of the rate, so the curve is lower. Isolating each task repeats the shared setup per task and does not grow with position. Illustrative sizes: a 5,000-token shared setup and 6,000-token outputs. items completed in the batch input billed (tokens) 0 1M 2M 1 10 20 30 One long session quadratic Same, with caching still quadratic, smaller Isolated per task setup repeats, no growth Illustrative: a 5,000-token shared setup and 6,000-token outputs, cache reads at a tenth of the input rate.

Found this useful? Share it:

Share on LinkedIn