Core AI Concepts Diagrams

Last updated:

Self-Attention

Structure

One token weighing every earlier token, heaviest on the noun it refers to.

Self-attention from one token to the tokens before it The sentence The animal didn't cross the street because it was tired, split into tokens. Curved lines run from the token it back to every earlier token, and a bar under each earlier token shows its attention weight. The heaviest line and tallest bar, 52 percent, go to animal, the noun it refers to. Street gets 18 percent, because 15 percent, and the rest a few percent each. The tokens after it, was and tired, are dashed out because a decoder-only model cannot see later tokens. The weights are illustrative, and a real model runs many attention heads. ATTENTION FROM “IT” TO EVERY EARLIER TOKEN The animal didn't cross the street because it was tired 3% 52% 4% 6% 2% 18% 15% the token being processed later tokens are masked in a decoder Thicker line, higher weight. Weights are illustrative, and a real model runs many attention heads, each weighting differently.

The Generation Loop

Flow

One pass through the model per token, each output appended to the input.

Token-by-token generation in a large language model The generation loop. The tokens so far, here The cat sat on the, go through one full pass of the transformer, which scores every token in the vocabulary. The scores become next-token probabilities: mat 41 percent, floor 22, couch 14, bed 9, roof 5, and a long tail of unlikely tokens. One token is sampled, with temperature and top-p shaping the pick. The sampled token is appended and the loop repeats, so the next pass reads The cat sat on the mat. The loop ends at a stop token or the output limit, and the tokens produced so far are the response. Tokens so far the prompt plus everything generated up to now The cat sat on the Transformer one full pass through every layer, scoring each token in the vocabulary Next-token probabilities mat 41% floor 22% couch 14% bed 9% roof 5% and a long tail of unlikely tokens Sample one token temperature and top-p shape which one is picked → “ mat” append the token and repeat next pass reads: The cat sat on the mat stop token or output limit Response Probabilities are illustrative.

Text to Tokens

Structure

Common words as single tokens, rare names and code split finer.

How text splits into tokens Three strings split into tokens, shown as alternating colored chips. The weather is nice today becomes five tokens, one per common word, about five characters each. Deploy to Kubernetes becomes five tokens because the rare name Kubernetes splits into Kub, ern, and etes. The JSON fragment with a retry field of 3 becomes five tokens from twelve characters, because punctuation and quotes take tokens of their own. The splits are illustrative and vary by tokenizer. Each piece maps to an integer ID in the vocabulary. TEXT TOKENS (ILLUSTRATIVE SPLIT) TOKENS Common words The weather is nice today The weather is nice today 5 5.0 chars each A rare name Deploy to Kubernetes Deploy to Kub ern etes 5 4.0 chars each Code and JSON {"retry": 3} {" retry ": 3 } 5 2.4 chars each Each piece maps to an integer ID in the model's vocabulary, which is all the model reads. Splits vary by tokenizer. Non-English text usually breaks into more pieces than English.

One Context Window, Shared

Layering

Prompt, history, documents, and output all drawn from one token budget.

Everything in a request shares one context window Two horizontal bars, each measured against a dashed context window limit. The first request holds a small system prompt, the conversation history, documents and tool results, a short new message, and room for the generated output and reasoning, with a little space to spare. In the second request the history and documents are larger and the input reaches the limit on its own, so the output segment sits outside the window with no room, and the request fails or the output is cut short. context window limit A request with room to answer conversation history documents, tool results output and reasoning Input and output are counted against the same limit. Input that fills the window conversation history documents, tool results output No room is left for the answer, so the request fails or the output is cut short. system prompt conversation history documents and tool results new message generated output

Every Request Resends the Conversation

Structure

Five chat requests, each resending all history, growing toward the limit.

Each chat request resends the whole conversation Five requests in one chat, drawn as horizontal bars. Each request sends the system prompt, every earlier user message and reply, and the new user message, then generates a new reply outside the input. Request 1 sends 450 tokens, request 2 sends 1,100, request 3 sends 1,750, request 4 sends 2,400, and request 5 sends 3,050, and its reply reaches the context window limit. Input billed across the five requests totals 8,750 tokens for a conversation of 3,050. Sizes are illustrative. context window limit Request 1 450 tokens in Request 2 1,100 tokens in Request 3 1,750 tokens in Request 4 2,400 tokens in Request 5 3,050 tokens in Input billed across the five requests: 8,750 tokens, for a conversation of 3,050. system prompt user message earlier reply, resent as input reply generated by this request Sizes are illustrative: a 300-token system prompt, 150-token messages, 500-token replies.

Temperature

Structure

One next-token distribution sharpened at low temperature, flattened at high.

Temperature rescaling a next-token distribution Three bar charts of the same six candidate tokens, mat, floor, couch, bed, roof, and moon, at three temperatures. At low temperature, 0.2, mat takes almost all the probability. At temperature 1.0 the model's own distribution is used, with mat at 45 percent and the others trailing. At high temperature, 1.8, the bars flatten and unlikely tokens such as roof and moon gain probability. Probabilities are illustrative. Low temperature (0.2) nearly always picks “mat” 95% mat 4% floor <1% couch <1% bed <1% roof <1% moon Temperature 1.0 the model's own distribution 45% mat 24% floor 15% couch 10% bed 5% roof 1% moon High temperature (1.8) unlikely tokens gain ground 32% mat 23% floor 18% couch 14% bed 10% roof 4% moon The same six candidates for the next token after “The cat sat on the”, rescaled at three temperatures. Probabilities are illustrative.

Top-p Sampling

Rules

The smallest set of likely tokens reaching p, with the tail cut off.

Top-p nucleus sampling cutting off unlikely tokens Ten candidate tokens sorted by probability, with a running sum under each: mat 41 percent, floor 22, couch 14, bed 9, roof 5, then rug, chair, table, moon, and sky at 3 percent or less. With top-p set to 0.9 a dashed line falls after roof, where the running sum reaches 91 percent. The five tokens left of the line are sampled from, rescaled to sum to 100 percent, and the five to the right are never sampled. Probabilities are illustrative. 41% mat 41% 22% floor 63% 14% couch 77% 9% bed 86% 5% roof 91% 3% rug 94% 2% chair 96% 2% table 98% 1% moon 99% 1% sky 100% top-p = 0.9 sampled from these 5, rescaled to sum to 100% never sampled TOKEN RUNNING SUM Top-p keeps the smallest set of most likely tokens whose probabilities reach p, here 91% after five tokens. Probabilities are illustrative.

Embeddings as Directions

Structure

Sentences as vectors, ranked by the angle between them and a query.

Sentence embeddings compared by cosine similarity Five sentences drawn as arrows of equal length from a shared origin, ending on the arc of a unit circle. How do I reset my password? and I forgot my login credentials point in nearly the same direction, 11 degrees apart, for a cosine similarity of 0.98. Change the account email sits 33 degrees away at 0.84. Forecast calls for rain and The weather is nice today point far away, at 0.39 and 0.17. A panel ranks the four sentences by cosine similarity to the password question. It is a two-dimensional sketch of vectors that really have hundreds to thousands of dimensions. How do I reset my password? I forgot my login credentials Change the account email Forecast calls for rain The weather is nice today COSINE SIMILARITY TO “How do I reset my password?” I forgot my login credentials 0.98 angle 11° Change the account email 0.84 angle 33° Forecast calls for rain 0.39 angle 67° The weather is nice today 0.17 angle 80° Smaller angle, closer meaning. A 2-D sketch: real embeddings have hundreds to thousands of dimensions, and normalized ones have length 1.

Mixture-of-Experts Routing

Flow

A router sending each token to 2 of 8 experts, all held in memory.

Mixture-of-experts routing inside one layer A mixture-of-experts layer. One token's representation goes to a router, which picks the top 2 of 8 experts. Experts 2 and 6 run, with weights 0.7 and 0.3, and the other six stay idle for this token. The two outputs are combined in a weighted sum and passed to the next layer. All eight experts sit in GPU memory. Compute follows the active parameters, 37 billion per token for DeepSeek-V3, while memory follows the total, all 671 billion. One token its representation at this layer Router picks the top 2 of 8 experts ALL EXPERTS IN GPU MEMORY Expert 1 idle Expert 2 weight 0.7 Expert 3 idle Expert 4 idle Expert 5 idle Expert 6 weight 0.3 Expert 7 idle Expert 8 idle Weighted sum passed to the next layer Compute follows active parameters DeepSeek-V3 runs 37B per token Memory follows total parameters DeepSeek-V3 loads all 671B

Found this useful? Share it:

Share on LinkedIn