The Multilayer Perceptron (MLP): A Complete Guide
Quick summary: A Multilayer Perceptron (MLP) is the classic feed-forward neural network: layers of simple computing units (neurons) stacked one after another, where every neuron in one layer connects to every neuron in the next. Trained with backpropagation and gradient descent, an MLP can learn to approximate remarkably complex functions. It is the foundation on which almost all of modern deep learning is built.
- What is an MLP?
- From the single neuron to the perceptron
- Architecture of an MLP
- Forward propagation (the maths)
- Activation functions
- Loss functions
- Backpropagation
- Optimization: gradient descent and variants
- Weight initialization
- Regularization and overfitting
- Hyperparameters
- Python example from scratch (NumPy)
- Strengths and limitations
- Applications
- Summary
1. What is an MLP?
A Multilayer Perceptron is a class of feed-forward artificial neural network. "Feed-forward" means information flows in one direction only: from the input, through one or more hidden layers, to the output, with no loops or cycles. "Multilayer" means there is at least one hidden layer between the input and the output. The word "perceptron" is historical: it comes from Frank Rosenblatt's 1958 perceptron, a single artificial neuron that could learn to separate two classes with a straight line.
A single perceptron can only solve linearly separable problems. The famous failure case is the XOR function, which Minsky and Papert pointed out in 1969. Stacking neurons into layers and adding non-linear activation functions removes this limitation. The MLP, trained with backpropagation (popularized by Rumelhart, Hinton and Williams in 1986), became the first widely successful multi-layer learner.
2. From the single neuron to the perceptron
An artificial neuron receives several inputs x1, …, xn, multiplies each by a weight wi, adds a bias b, and passes the result through an activation function f:
a = f(z)
- Weights (w): how strongly each input influences the neuron. These are learned.
- Bias (b): a learned offset that shifts the activation threshold.
- Pre-activation (z): the weighted sum before the non-linearity.
- Activation (a): the neuron's output after the non-linearity.
Geometrically, the equation w·x + b = 0 defines a hyperplane. A single neuron with a step activation classifies points by which side of the hyperplane they fall on. An MLP combines many such hyperplanes, bent by non-linear activations, to carve out arbitrarily complicated decision regions.
3. Architecture of an MLP
Figure: A fully connected MLP with 4 inputs, two hidden layers (4 and 3 neurons) and 2 outputs.
The three kinds of layers
| Layer | Role | Size is determined by |
|---|---|---|
| Input layer | Holds the raw feature vector. It performs no computation; it just passes the values on. | Number of input features (e.g., 784 for a 28×28 image). |
| Hidden layer(s) | Learn intermediate representations through weighted sums and non-linear activations. | A design choice (hyperparameter). |
| Output layer | Produces the final prediction. | Task: 1 unit for regression or binary classification, K units for K-class classification. |
Every layer is fully connected (also called dense): each neuron receives input from every neuron in the previous layer. A network with L layers having nl neurons in layer l has a weight matrix W(l) of shape nl × nl-1 and a bias vector b(l) of length nl.
Counting parameters. For a network 784 → 128 → 64 → 10:
Layer 2: 128×64 + 64 = 8,256
Layer 3: 64×10 + 10 = 650
Total = 109,386 trainable parameters
4. Forward propagation (the maths)
Forward propagation is the process of computing the network's output for a given input. Let a(0) = x be the input vector. For each layer l = 1, …, L:
a(l) = f(l)( z(l) )
Å· = a(L)
The whole network is therefore a composition of functions:
Worked mini-example
Take one input vector x = [1, 2], one hidden layer with 2 neurons (ReLU) and one output neuron (linear).
z(1) = [0.5·1 + (-1.0)·2 + 0.1, 0.3·1 + 0.8·2 - 0.2] = [-1.4, 1.7]
a(1) = ReLU(z(1)) = [0, 1.7]
W(2) = [2.0, 1.0], b(2) = 0.5
Å· = 2.0·0 + 1.0·1.7 + 0.5 = 2.2
5. Activation functions
| Function | Formula | Range | Notes |
|---|---|---|---|
| Sigmoid | σ(z) = 1 / (1 + e-z) | (0, 1) | Interpretable as a probability; used for binary output. Saturates, so gradients vanish for large |z|. Not zero-centered. |
| Tanh | tanh(z) = (ez - e-z) / (ez + e-z) | (-1, 1) | Zero-centered, so usually better than sigmoid in hidden layers, but still saturates. |
| ReLU | max(0, z) | [0, ∞) | The default for hidden layers. Cheap, does not saturate for z > 0. Can suffer "dead neurons" when z < 0 always. |
| Leaky ReLU | max(αz, z), α ≈ 0.01 | (-∞, ∞) | Fixes dead ReLUs by keeping a small slope for negative inputs. |
| GELU / SiLU (Swish) | z·Î¦(z) / z·Ïƒ(z) | ~(-0.28, ∞) | Smooth ReLU-like functions used in modern networks such as Transformers. |
| Softmax | ezi / Σj ezj | (0, 1), sums to 1 | Output layer for multi-class classification; produces a probability distribution. |
| Linear (identity) | f(z) = z | (-∞, ∞) | Output layer for regression. |
6. Loss functions
The loss (cost) function measures how far the prediction Å· is from the true target y. Training means finding weights that minimize its average over the training set.
| Task | Output activation | Loss |
|---|---|---|
| Regression | Linear | Mean Squared Error: L = (1/N) Σ (yi - ŷi)2 |
| Binary classification | Sigmoid | Binary cross-entropy: L = -(1/N) Σ [ y log ŷ + (1-y) log(1-ŷ) ] |
| Multi-class classification | Softmax | Categorical cross-entropy: L = -(1/N) Σi Σk yik log ŷik |
Cross-entropy is preferred for classification because, combined with sigmoid or softmax, its gradient with respect to the pre-activation simplifies to the neat expression (Å· - y), which avoids slow learning when the output saturates.
7. Backpropagation
To reduce the loss we need the gradient ∂L/∂W and ∂L/∂b for every layer. Backpropagation computes all of them efficiently by applying the chain rule of calculus backwards from the output layer to the input layer, re-using intermediate results rather than recomputing them.
The four key equations
Define the error signal of layer l as δ(l) = ∂L/∂z(l).
(2) Propagate back: δ(l) = ( (W(l+1))T δ(l+1) ) ⊙ f'(z(l))
(3) Weight gradient: ∂L/∂W(l) = δ(l) (a(l-1))T
(4) Bias gradient: ∂L/∂b(l) = δ(l)
(⊙ denotes element-wise multiplication.) For softmax with cross-entropy, equation (1) reduces to δ(L) = Å· - y.
Algorithm overview
- Forward pass: compute and store z(l) and a(l) for every layer.
- Compute the loss for the mini-batch.
- Backward pass: compute δ(L), then δ(L-1), …, δ(1) using equations (1) and (2), and the gradients with (3) and (4).
- Update the parameters with an optimizer (next section).
- Repeat for all mini-batches (one pass = one epoch) and for many epochs.
8. Optimization: gradient descent and variants
Given the gradient, parameters are updated in the direction that decreases the loss:
where η is the learning rate. Variants differ in how much data is used per step and how the step is computed:
- Batch gradient descent: uses the whole dataset per update. Stable but slow and memory heavy.
- Stochastic gradient descent (SGD): one sample per update. Fast and noisy; the noise can help escape poor minima.
- Mini-batch SGD: the practical standard (batch sizes of 32 to 512). Balances speed, stability and hardware efficiency.
- Momentum: accumulates a moving average of past gradients to smooth the path and accelerate along consistent directions.
- RMSProp / Adagrad: adapt the learning rate per parameter using the history of squared gradients.
- Adam / AdamW: combines momentum with adaptive learning rates; the most common default optimizer today.
Learning-rate scheduling (step decay, cosine annealing, warm-up) usually improves final accuracy. A learning rate that is too high makes the loss oscillate or diverge; too low makes training painfully slow.
9. Weight initialization
Initializing all weights to zero makes every neuron in a layer compute the same thing and receive the same gradient, so they never differentiate (the symmetry problem). Weights must be initialized randomly, with a scale chosen to keep activations and gradients at a stable magnitude across layers:
- Xavier / Glorot (for tanh, sigmoid): Var(W) = 2 / (nin + nout)
- He / Kaiming (for ReLU): Var(W) = 2 / nin
- Biases are commonly initialized to zero.
10. Regularization and overfitting
An MLP with many parameters can memorize the training set (overfitting): training loss keeps falling while validation loss rises. Common countermeasures:
| Technique | How it helps |
|---|---|
| L2 regularization (weight decay) | Adds λΣw2 to the loss, keeping weights small and the function smoother. |
| L1 regularization | Adds λΣ|w|, encouraging sparse weights. |
| Dropout | Randomly zeroes a fraction of neurons (e.g., 20-50%) during training so neurons cannot co-adapt; acts like averaging many sub-networks. |
| Early stopping | Stop training when validation loss stops improving. |
| Batch / Layer normalization | Normalizes layer inputs; stabilizes and speeds up training with a mild regularizing effect. |
| Data augmentation & more data | The most reliable cure for overfitting. |
| Smaller network | Fewer parameters reduce capacity to memorize. |
Universal Approximation Theorem
A theoretical result (Cybenko 1989, Hornik 1991) states that an MLP with a single hidden layer containing enough neurons and a non-linear (non-polynomial) activation can approximate any continuous function on a compact domain to arbitrary accuracy. The theorem guarantees that a good solution exists; it does not say how many neurons are needed or that gradient descent will find it. In practice, deeper networks represent many functions far more efficiently than very wide shallow ones.
11. Hyperparameters
| Hyperparameter | Typical choices / advice |
|---|---|
| Number of hidden layers | Start with 1 to 3; add depth only if validation performance improves. |
| Neurons per layer | Often 32 to 512; powers of two are common, funnel (decreasing) or constant widths both work. |
| Activation function | ReLU for hidden layers; task-specific output activation. |
| Learning rate | Most important setting. Try 1e-2 to 1e-4 with Adam; use a schedule. |
| Batch size | 32 to 256 is a good starting range. |
| Epochs | Use early stopping rather than a fixed guess. |
| Regularization strength | Dropout 0.1 to 0.5; weight decay 1e-5 to 1e-2. |
| Optimizer | Adam / AdamW as default; SGD with momentum when you want best generalization and can tune. |
Preprocessing matters: standardize numeric inputs (zero mean, unit variance) or scale them to [0, 1]; one-hot encode categorical features; always fit scalers on the training set only.
12. Python example from scratch (NumPy)
The following minimal MLP (one hidden layer, ReLU, sigmoid output) learns the XOR function, the problem a single perceptron cannot solve.
import numpy as np
# XOR dataset
X = np.array([[0,0],[0,1],[1,0],[1,1]], dtype=float)
y = np.array([[0],[1],[1],[0]], dtype=float)
rng = np.random.default_rng(42)
n_in, n_hidden, n_out = 2, 8, 1
# He initialization
W1 = rng.normal(0, np.sqrt(2/n_in), (n_in, n_hidden)); b1 = np.zeros((1, n_hidden))
W2 = rng.normal(0, np.sqrt(2/n_hidden), (n_hidden, n_out)); b2 = np.zeros((1, n_out))
relu = lambda z: np.maximum(0, z)
relu_d = lambda z: (z > 0).astype(float)
sigmoid = lambda z: 1 / (1 + np.exp(-z))
lr = 0.1
for epoch in range(5000):
# ---- forward pass ----
z1 = X @ W1 + b1
a1 = relu(z1)
z2 = a1 @ W2 + b2
y_hat = sigmoid(z2)
# ---- loss (binary cross-entropy) ----
eps = 1e-9
loss = -np.mean(y*np.log(y_hat+eps) + (1-y)*np.log(1-y_hat+eps))
# ---- backward pass ----
d2 = (y_hat - y) / len(X) # sigmoid + BCE simplification
dW2 = a1.T @ d2; db2 = d2.sum(axis=0, keepdims=True)
d1 = (d2 @ W2.T) * relu_d(z1)
dW1 = X.T @ d1; db1 = d1.sum(axis=0, keepdims=True)
# ---- gradient descent update ----
W2 -= lr * dW2; b2 -= lr * db2
W1 -= lr * dW1; b1 -= lr * db1
if epoch % 1000 == 0:
print(f"epoch {epoch:4d} loss {loss:.4f}")
print("Predictions:", y_hat.round(3).ravel()) # approx [0, 1, 1, 0]
The same model in PyTorch
import torch, torch.nn as nn
model = nn.Sequential(
nn.Linear(784, 128), nn.ReLU(), nn.Dropout(0.2),
nn.Linear(128, 64), nn.ReLU(),
nn.Linear(64, 10) # raw logits; CrossEntropyLoss applies softmax
)
loss_fn = nn.CrossEntropyLoss()
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
for xb, yb in train_loader:
opt.zero_grad()
loss = loss_fn(model(xb), yb)
loss.backward() # backpropagation
opt.step() # parameter update
13. Strengths and limitations
| Strengths | Limitations |
|---|---|
|
|
14. Applications
- Tabular prediction: credit scoring, churn prediction, house price estimation, demand forecasting.
- Classification: handwritten digits (MNIST), spam detection, medical diagnosis from clinical features.
- Function approximation and surrogate modelling in engineering and physics.
- Building blocks inside larger models: the feed-forward sub-layer in every Transformer block is an MLP; classification heads on top of CNN or language-model features are MLPs; the MLP-Mixer family replaces attention and convolution entirely with MLPs.
- Reinforcement learning: policy and value networks for low-dimensional state spaces.
15. Summary
- An MLP is a stack of fully connected layers: input → hidden layer(s) → output.
- Each layer computes a = f(Wx + b); non-linear activations are essential.
- Training minimizes a loss via backpropagation (chain rule) and gradient descent variants such as Adam.
- Good results depend on proper initialization, input scaling, regularization and sensible hyperparameters.
- Despite its simplicity, the MLP remains a core component of nearly every modern deep-learning architecture.
Found this useful? Share it, and leave a comment with your questions about neural networks.

0 Comments