# Generative Model Architectures


📂 model-architectures

## MODEL ARCHITECTURES

Model architecture field guide

# Generative Model Architectures

Understand what actually sits behind text, embedding, image, video, audio, and multimodal products, then choose a backbone, representation, generator, and system wrapper deliberately.

22 architecture families · 5 groups · one page per family

THE ESSENTIAL DISTINCTION

## A capability is not an architecture

“Chat,” “image,” and “multimodal” describe an interface. Decoder-only, state-space, DiT, and MoE describe internal computation. RAG and tool use describe the runtime system around a model. One production system can legitimately carry all three kinds of label.

## Write an architecture passport

- 1
Capability: Inputs, outputs, user job, and required evidence

- 2
Representation: Tokens, vectors, pixels, codec codes, or continuous latents

- 3
Backbone: Encoder, decoder, seq2seq, attention, SSM, or hybrid

- 4
Capacity: Dense or expert-routed; total versus active parameters

- 5
Generation: Autoregressive, diffusion/flow, adversarial, or deterministic

- 6
System: Retrieval, tools, validators, permissions, monitoring, and humans

## Browse all 22 families

Each family has its own page with training, inference, strengths, trade-offs, and selection guidance.

### PRODUCT CAPABILITIES

What the product accepts and returns. These are deployment shapes, not mutually exclusive neural architectures.

#### Text & chat language models

Generate text one token at a time, usually with a causal decoder trained for next-token prediction and then adapted to follow instructions.

#### Embedding models

Map text, images, audio, users, or items into vectors whose geometry is trained to preserve a useful notion of similarity.

#### Image generation & editing

Synthesize or transform pixels, most often through a text or image conditioner, a latent generator, and an image decoder.

#### Video generation

Generate spatial and temporal structure together using frame, patch, or latent video representations with cross-frame computation.

#### Audio, speech & music models

A family of task-specific pipelines: speech recognition, speech synthesis, audio understanding, voice conversion, sound generation, and music generation are not one architecture.

### CORE BACKBONES

How information moves through the trainable network. A deployed model may combine several of these choices.

#### Encoder-only Transformers

Let every input token attend to tokens on both sides, producing contextual representations rather than an open-ended generation loop.

#### Decoder-only Transformers

Use a causal mask so each position sees only earlier tokens, matching the left-to-right generation process used by most general-purpose chat models.

#### Encoder–decoder Transformers

Encode an input bidirectionally, then generate an output causally while cross-attending to the encoded source.

#### Mixture of Experts (MoE)

Replace selected dense sublayers with many expert networks and a learned router that activates only a small subset for each token.

#### State-space & recurrent models

Update a compact state as tokens arrive instead of retaining an explicit attention relationship between every pair of positions.

### REPRESENTATION & GENERATION

How images, audio, video, and other high-dimensional outputs are represented, learned, and sampled.

#### Autoencoders, VAEs & learned tokenizers

Learn an encoder that compresses data and a decoder that reconstructs it; variational and quantized variants shape the latent space for sampling or downstream generation.

#### Generative Adversarial Networks (GANs)

Train a generator to fool a discriminator while the discriminator learns to distinguish generated samples from training data.

#### Normalizing flows

Transform a simple distribution into a complex data distribution through a sequence of invertible mappings with tractable Jacobians.

#### Diffusion & score-based generation

Learn to reverse a gradual noising process, producing data by repeatedly transforming noise toward a sample.

#### Latent diffusion

Run diffusion in an autoencoder’s lower-dimensional latent space, then decode the generated representation back to pixels or another signal.

#### Diffusion Transformers (DiT)

Use a Transformer over noisy image or video patches as the denoising network inside a diffusion or flow-based generative process.

#### Autoregressive visual & audio models

Convert images, video, or audio into discrete codes and predict those codes in an ordering, often with a causal Transformer.

### MODALITY INTERFACES

How vision, text, audio, and other modalities are encoded, aligned, fused, or connected to a generator.

#### Vision encoders

Transform pixels into feature maps, patch tokens, or pooled vectors that downstream classifiers, retrievers, decoders, or multimodal models can consume.

#### Contrastive & dual encoders

Encode two inputs independently (such as a query and document or image and caption) and train matching pairs to land near each other.

#### Multimodal fusion models

Connect modality-specific encoders and generators through projection, cross-attention, shared token spaces, or combinations of these mechanisms.

### AI SYSTEM ARCHITECTURES

Runtime structures around a model that add knowledge, tools, control, and verifiable behavior.

#### Retrieval-augmented & hybrid systems

Retrieve evidence at request time and provide it to a generator, classifier, or agent instead of relying only on model parameters.

#### Tool-using & reasoning systems

Place a model inside a controller that can plan, call typed tools, inspect results, revise, verify, stop, and hand control to a person.

## Model Architectures

## Model Architectures
