Technical language is useful only when it makes the work clearer.
This lesson defines the vocabulary people need before they choose a local model, connect company information, or allow a system to act. The terms are organized as five layers rather than an alphabetical list: the learned core, the running moment, the information layer, the compute layer, and the action layer.
One branch-operations example connects the entire lesson. A system finds customer quotes with no recorded follow-up, retrieves the source record, drafts a next action, and stops for a manager's approval. That modest job makes the boundaries visible: the model is not the CRM, context is not memory, retrieval is not truth, a tool is not permission, and an agent is not unrestricted autonomy.
The goal is not to sound technical. It is to know where capability lives, what changes, what persists, what costs compute, and where authority should stop.
WHAT YOU WILL LEARN
- What a model and its learned parameters are - and what they are not.
- How tokens, context, and inference describe the moment a trained model is running.
- How embeddings, retrieval-augmented generation, and memory bring outside information into useful work.
- What quantization changes, and how GPUs, VRAM, and unified memory affect what can run locally.
- How APIs, tools, workflows, agents, and approval gates form the action layer around a model.
- Why precise language prevents expensive category errors in AI purchasing and implementation.