Core AI Concepts Diagrams
Self-Attention
StructureOne token weighing every earlier token, heaviest on the noun it refers to.
The Generation Loop
FlowOne pass through the model per token, each output appended to the input.
Text to Tokens
StructureCommon words as single tokens, rare names and code split finer.
One Context Window, Shared
LayeringPrompt, history, documents, and output all drawn from one token budget.
Every Request Resends the Conversation
StructureFive chat requests, each resending all history, growing toward the limit.
Temperature
StructureOne next-token distribution sharpened at low temperature, flattened at high.
Top-p Sampling
RulesThe smallest set of likely tokens reaching p, with the tail cut off.
Embeddings as Directions
StructureSentences as vectors, ranked by the angle between them and a query.
Mixture-of-Experts Routing
FlowA router sending each token to 2 of 8 experts, all held in memory.
Found this useful? Share it:
Share on LinkedIn