Glossary · 1 · Foundations: how AI works
Transformer
The transformer is the neural network architecture behind today’s large language models. It uses a mechanism called attention to weigh how strongly every token in a sequence relates to every other token, which lets it model long-range context and train efficiently on very large amounts of text.
- Intermediate
- Expert
- Developers
In one sentence
The transformer architecture explained: attention over tokens, the design behind GPT, Claude, Gemini, Llama and nearly every modern LLM.
Example
In “The pump stops when the valve closes because it senses the pressure drop,” attention helps the model link “it” to “the pump” rather than “the valve.”
Why it matters on your learning path
- Technical writers: Understanding attention explains why clear, unambiguous references in your source content help AI answers.
- Technical project managers: Transformer-based models dominate the market; architecture is rarely a selection criterion, model quality and terms of use are.
- Developers: Knowing the architecture helps you reason about context length, cost per token and why output is generated one token at a time.
Background
The architecture was introduced by Google researchers in the 2017 paper “Attention Is All You Need.” The “T” in GPT stands for transformer.