Engineering2026-09-172 min read

decoder-only-transformers

VDaily Team
Maintainer

Decoder-only Transformers

📂 model-architectures

MODEL ARCHITECTURES#

Decoder-only Transformers

Use a causal mask so each position sees only earlier tokens, matching the left-to-right generation process used by most general-purpose chat models.

MENTAL MODEL#

Compress the prefix into attention state, predict one continuation token, append it, and repeat.

DATA FLOW#

  • Prefix tokens

  • Causally masked self-attention

  • Hidden state at each position

  • Vocabulary logits

  • Next token → repeat

How it trains#

Next-token cross-entropy supplies a learning signal at every position. Post-training may add instruction examples, preferences, verifiable rewards, or tool trajectories.

How inference runs#

Prefill processes the prompt; decode then emits tokens serially while the KV cache stores past attention keys and values. Context size, cache memory, and output length drive serving behavior.

Strengths#

  • • One scalable interface for many generative tasks

  • • Strong in-context adaptation

  • • Can interleave prose, code, structured data, and tool-call tokens

Trade-offs#

  • • Serial decode latency

  • • Quadratic prefill in standard full attention

  • • Prompt sensitivity and probabilistic output require validation

Use it when#

  • Open-ended generation or interactive chat is central

  • Many tasks can share one model endpoint

  • You can constrain and evaluate outputs in the product loop

Avoid or challenge it when#

  • A small encoder or deterministic component meets the requirement

  • Guaranteed exact outputs are assumed without validation

  • Long source corpora should be searched rather than placed wholesale in context

Illustrative published families#

  • • GPT-style causal language models

  • • LLaMA research family

Commonly combines with#

PRIMARY REFERENCES#

Model Architectures#

Model Architectures#

Tags:
decoder-only-transformers — Blog — VDaily