autoencoders-vaes-learned-tokenizers
Autoencoders, VAEs & learned tokenizers
📂 model-architectures
MODEL ARCHITECTURES#
Autoencoders, VAEs & learned tokenizers
Learn an encoder that compresses data and a decoder that reconstructs it; variational and quantized variants shape the latent space for sampling or downstream generation.
MENTAL MODEL#
A learned codec. It decides which information survives before another model generates in the compressed space.
DATA FLOW#
-
Image, audio, or other signal
-
Encoder
-
Continuous latent or discrete codes
-
Optional latent generator
-
Decoder reconstruction
How it trains#
A reconstruction objective preserves perceptual content. VAEs add a distribution regularizer; vector-quantized models map representations to a discrete codebook and add codebook/commitment objectives.
How inference runs#
Encode for compression or editing; sample or generate a latent/code sequence; decode to the original modality. Reconstruction quality limits the best possible downstream output.
Strengths#
-
• Reduces the dimensionality of expensive media
-
• Creates reusable continuous latents or discrete tokens
-
• Separates representation learning from the generative prior
Trade-offs#
-
• Compression discards detail
-
• Latent geometry may not match downstream semantics
-
• Quantized codebooks can be underused or introduce artifacts
Use it when#
-
Raw media is too expensive to model directly
-
A generative model needs a compact continuous or discrete space
-
Reconstruction failures are measured by content slice
Avoid or challenge it when#
-
Lossless reconstruction is mandatory
-
The codec domain differs materially from production data
-
Decoder artifacts would be mistaken for generator failures
Illustrative published families#
-
• Variational Autoencoder (VAE)
-
• VQ-VAE discrete visual/audio tokenizers
-
• Autoencoders used by latent diffusion