Deep learning has moved from a research curiosity to the engine behind search, translation, voice assistants, medical imaging, self-driving systems and generative AI. Behind all of this sits a small family of core architectures. Each one is a different way of wiring neurons together so that the network matches the structure of the data it is learning from: grids of pixels, sequences of words, graphs of relationships, and so on.
This guide walks through the most important architectures: how they work, what problem each one solves, where it breaks down, and when to use it.
- Feedforward Networks (MLP)
- Convolutional Neural Networks (CNN)
- Recurrent Neural Networks (RNN)
- LSTM and GRU
- Sequence-to-Sequence and Attention
- Transformers
- Autoencoders and VAEs
- Generative Adversarial Networks (GAN)
- Diffusion Models
- Graph Neural Networks (GNN)
- Mixture of Experts (MoE)
- Comparison Table
- How to Choose an Architecture
The Building Blocks Shared by All Architectures
Before the architectures themselves, it helps to know the pieces they are all built from:
- Neuron (unit): computes a weighted sum of its inputs plus a bias, then applies an activation function.
- Activation function: adds non-linearity. Without it, a stack of layers collapses into a single linear function. Common choices are ReLU, GELU, sigmoid, tanh and softmax.
- Loss function: measures how wrong the prediction is, for example cross-entropy for classification or mean squared error for regression.
- Backpropagation and gradient descent: the loss is differentiated with respect to every weight, and the weights are nudged in the direction that reduces it. Optimizers such as SGD with momentum and Adam do this nudging.
- Regularization: dropout, weight decay, data augmentation and early stopping keep the model from memorizing the training set.
- Normalization: batch norm and layer norm keep activations well scaled, which makes deep networks trainable.
1. Feedforward Networks (MLP)
Tabular dataFoundationThe multilayer perceptron is the original deep network. Data flows in one direction, from input layer through one or more hidden layers to the output layer. Every neuron in a layer connects to every neuron in the next, which is why these are also called fully connected or dense layers.
Strengths
- Simple, fast, and a universal function approximator in theory.
- A solid baseline for tabular and structured data.
- Appears as a sub-component inside nearly every other architecture, such as the feed-forward block in a Transformer.
Limitations
- Ignores structure: an image flattened into a vector loses its spatial layout.
- Parameter count explodes with input size.
- No memory of earlier inputs.
2. Convolutional Neural Networks (CNN)
ImagesVideoAudioCNNs are designed for data laid out on a grid. Instead of connecting every input to every neuron, a CNN slides small learnable filters (kernels) across the input. Each filter detects a local pattern such as an edge, a texture or a corner, and the same filter is reused at every position. This idea is called weight sharing, and it is why CNNs need far fewer parameters than dense networks on images.
Key components
- Convolution layer: applies filters to produce feature maps.
- Pooling layer: downsamples (max or average pooling) so the network becomes tolerant to small shifts and the computation shrinks.
- Stride and padding: control how far the filter moves and what happens at the borders.
- Hierarchy of features: early layers learn edges, middle layers learn shapes and parts, deep layers learn whole objects.
Landmark CNNs
| Model | Main idea |
|---|---|
| LeNet | Early CNN for handwritten digit recognition. |
| AlexNet | Showed deep CNNs on GPUs could win large-scale image recognition; used ReLU and dropout. |
| VGG | Very deep stacks of small 3x3 filters. |
| ResNet | Skip (residual) connections let networks go hundreds of layers deep without vanishing gradients. |
| Inception | Parallel filters of several sizes inside one block. |
| EfficientNet | Scales depth, width and resolution together in a balanced way. |
3. Recurrent Neural Networks (RNN)
SequencesTime seriesAn RNN processes a sequence one element at a time while carrying a hidden state, a running summary of everything seen so far. At each step the new hidden state is computed from the current input and the previous hidden state, using the same weights at every step.
Strengths
- Handles variable-length input naturally.
- Has a built-in notion of order and memory.
Limitations
- Vanishing and exploding gradients: during backpropagation through time, gradients are multiplied many times, so they shrink toward zero or blow up. Plain RNNs therefore struggle with long-range dependencies.
- Sequential computation, so training cannot be parallelized across time steps.
Variants include bidirectional RNNs, which read the sequence forwards and backwards, and stacked (deep) RNNs.
4. LSTM and GRU
Long memoryGatedLong Short-Term Memory networks were built to fix the vanishing-gradient problem. An LSTM keeps a separate cell state, a kind of conveyor belt that information can travel along with little change, and uses gates to control what enters, stays and leaves.
- Forget gate: decides what to erase from the cell state.
- Input gate: decides what new information to write.
- Output gate: decides what part of the cell state to expose as the hidden state.
The GRU (Gated Recurrent Unit) is a lighter cousin. It merges the cell and hidden state and uses just two gates (update and reset), giving fewer parameters and often similar accuracy.
5. Sequence-to-Sequence and Attention
TranslationBridge to TransformersA sequence-to-sequence (seq2seq) model has two parts: an encoder that reads the input sequence and compresses it into a vector, and a decoder that generates the output sequence from that vector. It made neural machine translation practical.
The weakness was the bottleneck: a whole sentence had to be squeezed into one fixed-size vector. Attention solved this. At every decoding step, the decoder computes a relevance score between its current state and every encoder position, turns those scores into weights with softmax, and takes a weighted average of the encoder states. In effect, the model learns where to look.
Attention proved so effective that researchers asked whether recurrence was needed at all. The answer was the Transformer.
6. Transformers
LanguageVisionMultimodalIntroduced in 2017 in the paper "Attention Is All You Need", the Transformer drops recurrence and convolution entirely and relies on self-attention. Every token in the input looks at every other token in a single step and decides how much each one matters. Because there is no step-by-step dependency, training parallelizes extremely well on GPUs and TPUs, which is what allowed models to scale to billions of parameters.
Core components
- Self-attention (queries, keys, values): each token produces a query, a key and a value vector. Attention weights come from comparing queries with keys, and the output is a weighted sum of values.
- Multi-head attention: several attention operations run in parallel, each free to focus on a different kind of relationship (syntax, coreference, position and so on).
- Positional encoding: since attention has no built-in notion of order, position information is added to the embeddings. Modern models often use rotary or relative position schemes.
- Feed-forward block: a small MLP applied to each token after attention.
- Residual connections and layer normalization: keep very deep stacks stable.
Three common layouts
| Layout | How it works | Examples | Best for |
|---|---|---|---|
| Encoder-only | Bidirectional attention; each token sees the whole input. | BERT, RoBERTa | Classification, search, embeddings |
| Decoder-only | Causal attention; each token sees only earlier tokens, so it can generate text. | The GPT family, Llama, Claude | Text and code generation, chat |
| Encoder-decoder | Encoder reads the input, decoder generates output while attending to it. | T5, original Transformer | Translation, summarization |
Beyond text
Vision Transformers (ViT) cut an image into patches and treat each patch like a token. Transformers also power speech models, protein structure prediction, and multimodal systems that combine text, images and audio.
Limitations
- Attention cost grows quadratically with sequence length, so long contexts are expensive. Techniques such as sparse attention, FlashAttention and KV caching reduce the pain.
- Large models need large data and large compute.
7. Autoencoders and VAEs
UnsupervisedRepresentation learningAn autoencoder learns to copy its input to its output through a narrow middle layer called the bottleneck or latent space. The encoder compresses the input, the decoder reconstructs it, and training minimizes the reconstruction error. Since the bottleneck cannot hold everything, the network is forced to learn the most useful features.
Variants
- Denoising autoencoder: trained to reconstruct a clean input from a corrupted one.
- Sparse autoencoder: encourages only a few latent units to be active.
- Variational autoencoder (VAE): the encoder outputs a probability distribution instead of a single point, and a KL-divergence term keeps the latent space smooth. You can sample from that space to generate new data.
8. Generative Adversarial Networks (GAN)
GenerativeAdversarialA GAN pits two networks against each other. The generator turns random noise into fake samples, and the discriminator tries to tell real samples from fake ones. As the discriminator gets better, the generator is pushed to produce more convincing output, and the cycle continues until the fakes are hard to distinguish from real data.
Notable variants
- DCGAN: convolutional GAN that made image generation stable.
- Conditional GAN: generation guided by a label or input.
- CycleGAN: unpaired image-to-image translation, such as horse to zebra.
- StyleGAN: high-quality, controllable face and image synthesis.
Limitations
- Training can be unstable.
- Mode collapse: the generator may produce only a narrow range of outputs.
- Hard to evaluate and to cover the full diversity of the data.
9. Diffusion Models
Image generationState of the artDiffusion models generate data by learning to reverse a noising process. In the forward direction, noise is added to a real image step by step until only static remains. A neural network, typically a U-Net or a Transformer, is trained to predict and remove the noise at each step. To generate something new, you start from pure noise and apply the denoiser repeatedly until a clean sample appears. Text prompts can steer the process through conditioning.
Why they took over
- Training is stable compared with GANs.
- They cover the diversity of the data well and produce high-fidelity results.
- Conditioning makes text-to-image, inpainting and editing natural.
Limitations
- Sampling needs many steps, so it is slower than a GAN's single pass. Faster samplers and distillation reduce this.
- Latent diffusion runs the process in a compressed latent space (using a VAE) to cut the cost, which is how many popular image generators work.
10. Graph Neural Networks (GNN)
GraphsRelational dataMany real-world datasets are not grids or sequences but graphs: social networks, molecules, road maps, knowledge bases. GNNs operate directly on nodes and edges using message passing. In each layer, every node collects information from its neighbours, combines it, and updates its own representation. After k layers, a node's representation reflects its k-hop neighbourhood.
- GCN (Graph Convolutional Network): averages neighbour features with learned weights.
- GraphSAGE: samples and aggregates neighbours, so it scales to large graphs.
- GAT (Graph Attention Network): uses attention to weigh neighbours differently.
Limitations
- Over-smoothing: with too many layers, node representations become nearly identical.
- Scaling to very large graphs needs sampling tricks.
11. Mixture of Experts (MoE)
ScalingEfficiencyA Mixture of Experts layer replaces one big feed-forward block with many smaller "expert" networks and a router (gating network). For each token, the router picks only a few experts, so the model can hold a very large number of parameters while using only a fraction of them per token. This decouples total capacity from compute per token, and it is widely used in large language models.
Trade-offs
- More capacity for the same compute per token.
- Harder to train: load balancing across experts matters, and memory use is high because all experts must be stored.
Comparison Table
| Architecture | Data type | Key idea | Main strength | Main weakness |
|---|---|---|---|---|
| MLP | Tabular, vectors | Dense layers | Simple, fast baseline | Ignores structure |
| CNN | Images, grids | Shared local filters | Efficient spatial features | Limited global context |
| RNN | Sequences | Recurrent hidden state | Natural order handling | Vanishing gradients, slow |
| LSTM / GRU | Sequences | Gated memory | Longer dependencies | Still sequential |
| Transformer | Text, images, audio, multimodal | Self-attention | Parallel, scales extremely well | Quadratic attention cost |
| Autoencoder / VAE | Any | Compress and reconstruct | Unsupervised features | Blurry generation (VAE) |
| GAN | Images, audio | Generator vs discriminator | Fast, sharp samples | Unstable, mode collapse |
| Diffusion | Images, video, audio | Learned denoising | Quality and diversity | Slow sampling |
| GNN | Graphs | Message passing | Uses relational structure | Over-smoothing, scaling |
| MoE | Usually inside Transformers | Sparse expert routing | Huge capacity, low per-token compute | Complex training, high memory |
How to Choose an Architecture
- Start from the structure of your data. Grid: CNN or ViT. Sequence or text: Transformer. Graph: GNN. Plain table: MLP or gradient-boosted trees.
- Consider the amount of data. Small datasets favour architectures with strong built-in assumptions (CNNs, GRUs) or transfer learning from a pretrained model.
- Consider the task. Understanding or classifying: encoder models. Generating: decoder-only Transformers or diffusion. Anomaly detection: autoencoders.
- Consider your compute and latency budget. Small CNNs and GRUs run well on phones and edge devices; large Transformers and diffusion models need serious hardware.
- Prefer pretrained models. For most practical work, fine-tuning an existing model beats training from scratch.
Conclusion
Deep learning architectures are best understood as different answers to one question: what structure does my data have, and how can the network exploit it? MLPs treat inputs as flat vectors, CNNs exploit spatial locality, RNNs and LSTMs exploit order, Transformers let every element attend to every other, GNNs follow edges, and generative models such as VAEs, GANs and diffusion models learn to create new data. Modern systems frequently combine several of these, for example a Transformer inside a diffusion model, or a CNN encoder feeding a Transformer decoder. Learn the core ideas here and new architectures will look like variations on a familiar theme.
Found this useful? Share it, and let me know in the comments which architecture you would like covered in depth next.

0 Comments