Engineering2026-09-172 min read

encoder-only-transformers

VDaily Team
Maintainer

Encoder-only Transformers

📂 model-architectures

MODEL ARCHITECTURES#

Encoder-only Transformers

Let every input token attend to tokens on both sides, producing contextual representations rather than an open-ended generation loop.

MENTAL MODEL#

Read the entire input, contextualize each position, then pool or attach a task head.

DATA FLOW#

  • Token sequence

  • Bidirectional self-attention blocks

  • Contextual token states

  • Pooler or task head

  • Class, span, score, or vector

How it trains#

Masked-token reconstruction is the classic pretraining objective, followed by task-specific classification, token labeling, span prediction, or metric learning.

How inference runs#

One forward pass processes the available input. It is naturally suited to understanding and representation, but does not by itself generate an arbitrary continuation.

Strengths#

  • • Rich bidirectional context for classification and extraction

  • • Parallel processing across the input

  • • Often efficient for high-volume bounded tasks

Trade-offs#

  • • Not an open-ended generator

  • • Maximum input length and quadratic attention can constrain long documents

  • • A task head or pooling strategy is usually required

Use it when#

  • Classification, named entities, reranking, or embeddings dominate

  • The complete input is available before prediction

  • A smaller specialized model can meet the rubric

Avoid or challenge it when#

  • The primary output is long free-form text

  • Streaming generation is required

  • Inputs routinely exceed the evaluated context strategy

Illustrative published families#

  • • BERT and RoBERTa research families

  • • Encoder-based rerankers and classifiers

Commonly combines with#

PRIMARY REFERENCES#

Model Architectures#

Model Architectures#

Tags:
encoder-only-transformers — Blog — VDaily