The Multilayer Perceptron (MLP)

multilayer perception_

The Multilayer Perceptron (MLP): A Complete Guide

Quick summary: A Multilayer Perceptron (MLP) is the classic feed-forward neural network: layers of simple computing units (neurons) stacked one after another, where every neuron in one layer connects to every neuron in the next. Trained with backpropagation and gradient descent, an MLP can learn to approximate remarkably complex functions. It is the foundation on which almost all of modern deep learning is built.

1. What is an MLP?

A Multilayer Perceptron is a class of feed-forward artificial neural network. "Feed-forward" means information flows in one direction only: from the input, through one or more hidden layers, to the output, with no loops or cycles. "Multilayer" means there is at least one hidden layer between the input and the output. The word "perceptron" is historical: it comes from Frank Rosenblatt's 1958 perceptron, a single artificial neuron that could learn to separate two classes with a straight line.

A single perceptron can only solve linearly separable problems. The famous failure case is the XOR function, which Minsky and Papert pointed out in 1969. Stacking neurons into layers and adding non-linear activation functions removes this limitation. The MLP, trained with backpropagation (popularized by Rumelhart, Hinton and Williams in 1986), became the first widely successful multi-layer learner.

Key idea: Each layer transforms its input into a new representation. Early layers learn simple features; deeper layers combine them into more abstract ones. The final layer turns the learned representation into a prediction.

2. From the single neuron to the perceptron

An artificial neuron receives several inputs x1, …, xn, multiplies each by a weight wi, adds a bias b, and passes the result through an activation function f:

z = w1x1 + w2x2 + … + wnxn + b = w·x + b
a = f(z)
  • Weights (w): how strongly each input influences the neuron. These are learned.
  • Bias (b): a learned offset that shifts the activation threshold.
  • Pre-activation (z): the weighted sum before the non-linearity.
  • Activation (a): the neuron's output after the non-linearity.

Geometrically, the equation w·x + b = 0 defines a hyperplane. A single neuron with a step activation classifies points by which side of the hyperplane they fall on. An MLP combines many such hyperplanes, bent by non-linear activations, to carve out arbitrarily complicated decision regions.

3. Architecture of an MLP

Input layer Hidden layer 1 Hidden layer 2 Output layer

Figure: A fully connected MLP with 4 inputs, two hidden layers (4 and 3 neurons) and 2 outputs.

The three kinds of layers

LayerRoleSize is determined by
Input layerHolds the raw feature vector. It performs no computation; it just passes the values on.Number of input features (e.g., 784 for a 28×28 image).
Hidden layer(s)Learn intermediate representations through weighted sums and non-linear activations.A design choice (hyperparameter).
Output layerProduces the final prediction.Task: 1 unit for regression or binary classification, K units for K-class classification.

Every layer is fully connected (also called dense): each neuron receives input from every neuron in the previous layer. A network with L layers having nl neurons in layer l has a weight matrix W(l) of shape nl × nl-1 and a bias vector b(l) of length nl.

Counting parameters. For a network 784 → 128 → 64 → 10:

Layer 1: 784×128 + 128 = 100,480
Layer 2: 128×64 + 64 = 8,256
Layer 3: 64×10 + 10 = 650
Total = 109,386 trainable parameters

4. Forward propagation (the maths)

Forward propagation is the process of computing the network's output for a given input. Let a(0) = x be the input vector. For each layer l = 1, …, L:

z(l) = W(l) a(l-1) + b(l)
a(l) = f(l)( z(l) )
Å· = a(L)

The whole network is therefore a composition of functions:

Å· = f(L)( W(L) f(L-1)( … f(1)( W(1)x + b(1) ) … ) + b(L) )

Worked mini-example

Take one input vector x = [1, 2], one hidden layer with 2 neurons (ReLU) and one output neuron (linear).

W(1) = [[0.5, -1.0], [0.3, 0.8]],   b(1) = [0.1, -0.2]
z(1) = [0.5·1 + (-1.0)·2 + 0.1,  0.3·1 + 0.8·2 - 0.2] = [-1.4, 1.7]
a(1) = ReLU(z(1)) = [0, 1.7]
W(2) = [2.0, 1.0],   b(2) = 0.5
Å· = 2.0·0 + 1.0·1.7 + 0.5 = 2.2
Why non-linearity matters: If every activation were linear, the composition of layers would collapse into a single linear transformation, Wtotalx + btotal. No matter how deep, the network would be no more powerful than linear regression. Non-linear activations are what give depth its power.

5. Activation functions

FunctionFormulaRangeNotes
Sigmoidσ(z) = 1 / (1 + e-z)(0, 1)Interpretable as a probability; used for binary output. Saturates, so gradients vanish for large |z|. Not zero-centered.
Tanhtanh(z) = (ez - e-z) / (ez + e-z)(-1, 1)Zero-centered, so usually better than sigmoid in hidden layers, but still saturates.
ReLUmax(0, z)[0, ∞)The default for hidden layers. Cheap, does not saturate for z > 0. Can suffer "dead neurons" when z < 0 always.
Leaky ReLUmax(αz, z), α ≈ 0.01(-∞, ∞)Fixes dead ReLUs by keeping a small slope for negative inputs.
GELU / SiLU (Swish)z·Î¦(z)  /  z·Ïƒ(z)~(-0.28, ∞)Smooth ReLU-like functions used in modern networks such as Transformers.
Softmaxezi / Σj ezj(0, 1), sums to 1Output layer for multi-class classification; produces a probability distribution.
Linear (identity)f(z) = z(-∞, ∞)Output layer for regression.
Rule of thumb: Hidden layers → ReLU (or a variant). Output layer → linear for regression, sigmoid for binary classification, softmax for multi-class classification.

6. Loss functions

The loss (cost) function measures how far the prediction Å· is from the true target y. Training means finding weights that minimize its average over the training set.

TaskOutput activationLoss
RegressionLinearMean Squared Error: L = (1/N) Σ (yi - ŷi)2
Binary classificationSigmoidBinary cross-entropy: L = -(1/N) Σ [ y log ŷ + (1-y) log(1-ŷ) ]
Multi-class classificationSoftmaxCategorical cross-entropy: L = -(1/N) Σi Σk yik log ŷik

Cross-entropy is preferred for classification because, combined with sigmoid or softmax, its gradient with respect to the pre-activation simplifies to the neat expression (Å· - y), which avoids slow learning when the output saturates.

7. Backpropagation

To reduce the loss we need the gradient ∂L/∂W and ∂L/∂b for every layer. Backpropagation computes all of them efficiently by applying the chain rule of calculus backwards from the output layer to the input layer, re-using intermediate results rather than recomputing them.

The four key equations

Define the error signal of layer l as δ(l) = ∂L/∂z(l).

(1) Output layer:   δ(L) = ∂L/∂a(L) ⊙ f'(z(L))
(2) Propagate back:   δ(l) = ( (W(l+1))T δ(l+1) ) ⊙ f'(z(l))
(3) Weight gradient:   ∂L/∂W(l) = δ(l) (a(l-1))T
(4) Bias gradient:   ∂L/∂b(l) = δ(l)

(⊙ denotes element-wise multiplication.) For softmax with cross-entropy, equation (1) reduces to δ(L) = Å· - y.

Algorithm overview

  1. Forward pass: compute and store z(l) and a(l) for every layer.
  2. Compute the loss for the mini-batch.
  3. Backward pass: compute δ(L), then δ(L-1), …, δ(1) using equations (1) and (2), and the gradients with (3) and (4).
  4. Update the parameters with an optimizer (next section).
  5. Repeat for all mini-batches (one pass = one epoch) and for many epochs.
Vanishing / exploding gradients: Equation (2) multiplies by WT and f' at every layer. If these factors are consistently smaller than 1 (e.g., sigmoid derivative ≤ 0.25), gradients shrink exponentially toward early layers; if larger than 1, they blow up. ReLU activations, careful initialization, normalization layers and residual connections are the standard remedies.

8. Optimization: gradient descent and variants

Given the gradient, parameters are updated in the direction that decreases the loss:

θ ← θ - η · ∇θL

where η is the learning rate. Variants differ in how much data is used per step and how the step is computed:

  • Batch gradient descent: uses the whole dataset per update. Stable but slow and memory heavy.
  • Stochastic gradient descent (SGD): one sample per update. Fast and noisy; the noise can help escape poor minima.
  • Mini-batch SGD: the practical standard (batch sizes of 32 to 512). Balances speed, stability and hardware efficiency.
  • Momentum: accumulates a moving average of past gradients to smooth the path and accelerate along consistent directions.
  • RMSProp / Adagrad: adapt the learning rate per parameter using the history of squared gradients.
  • Adam / AdamW: combines momentum with adaptive learning rates; the most common default optimizer today.

Learning-rate scheduling (step decay, cosine annealing, warm-up) usually improves final accuracy. A learning rate that is too high makes the loss oscillate or diverge; too low makes training painfully slow.

9. Weight initialization

Initializing all weights to zero makes every neuron in a layer compute the same thing and receive the same gradient, so they never differentiate (the symmetry problem). Weights must be initialized randomly, with a scale chosen to keep activations and gradients at a stable magnitude across layers:

  • Xavier / Glorot (for tanh, sigmoid): Var(W) = 2 / (nin + nout)
  • He / Kaiming (for ReLU): Var(W) = 2 / nin
  • Biases are commonly initialized to zero.

10. Regularization and overfitting

An MLP with many parameters can memorize the training set (overfitting): training loss keeps falling while validation loss rises. Common countermeasures:

TechniqueHow it helps
L2 regularization (weight decay)Adds λΣw2 to the loss, keeping weights small and the function smoother.
L1 regularizationAdds λΣ|w|, encouraging sparse weights.
DropoutRandomly zeroes a fraction of neurons (e.g., 20-50%) during training so neurons cannot co-adapt; acts like averaging many sub-networks.
Early stoppingStop training when validation loss stops improving.
Batch / Layer normalizationNormalizes layer inputs; stabilizes and speeds up training with a mild regularizing effect.
Data augmentation & more dataThe most reliable cure for overfitting.
Smaller networkFewer parameters reduce capacity to memorize.

Universal Approximation Theorem

A theoretical result (Cybenko 1989, Hornik 1991) states that an MLP with a single hidden layer containing enough neurons and a non-linear (non-polynomial) activation can approximate any continuous function on a compact domain to arbitrary accuracy. The theorem guarantees that a good solution exists; it does not say how many neurons are needed or that gradient descent will find it. In practice, deeper networks represent many functions far more efficiently than very wide shallow ones.

11. Hyperparameters

HyperparameterTypical choices / advice
Number of hidden layersStart with 1 to 3; add depth only if validation performance improves.
Neurons per layerOften 32 to 512; powers of two are common, funnel (decreasing) or constant widths both work.
Activation functionReLU for hidden layers; task-specific output activation.
Learning rateMost important setting. Try 1e-2 to 1e-4 with Adam; use a schedule.
Batch size32 to 256 is a good starting range.
EpochsUse early stopping rather than a fixed guess.
Regularization strengthDropout 0.1 to 0.5; weight decay 1e-5 to 1e-2.
OptimizerAdam / AdamW as default; SGD with momentum when you want best generalization and can tune.

Preprocessing matters: standardize numeric inputs (zero mean, unit variance) or scale them to [0, 1]; one-hot encode categorical features; always fit scalers on the training set only.

12. Python example from scratch (NumPy)

The following minimal MLP (one hidden layer, ReLU, sigmoid output) learns the XOR function, the problem a single perceptron cannot solve.

import numpy as np

# XOR dataset
X = np.array([[0,0],[0,1],[1,0],[1,1]], dtype=float)
y = np.array([[0],[1],[1],[0]], dtype=float)

rng = np.random.default_rng(42)
n_in, n_hidden, n_out = 2, 8, 1

# He initialization
W1 = rng.normal(0, np.sqrt(2/n_in), (n_in, n_hidden));  b1 = np.zeros((1, n_hidden))
W2 = rng.normal(0, np.sqrt(2/n_hidden), (n_hidden, n_out)); b2 = np.zeros((1, n_out))

relu     = lambda z: np.maximum(0, z)
relu_d   = lambda z: (z > 0).astype(float)
sigmoid  = lambda z: 1 / (1 + np.exp(-z))

lr = 0.1
for epoch in range(5000):
    # ---- forward pass ----
    z1 = X @ W1 + b1
    a1 = relu(z1)
    z2 = a1 @ W2 + b2
    y_hat = sigmoid(z2)

    # ---- loss (binary cross-entropy) ----
    eps = 1e-9
    loss = -np.mean(y*np.log(y_hat+eps) + (1-y)*np.log(1-y_hat+eps))

    # ---- backward pass ----
    d2 = (y_hat - y) / len(X)            # sigmoid + BCE simplification
    dW2 = a1.T @ d2;   db2 = d2.sum(axis=0, keepdims=True)
    d1 = (d2 @ W2.T) * relu_d(z1)
    dW1 = X.T @ d1;    db1 = d1.sum(axis=0, keepdims=True)

    # ---- gradient descent update ----
    W2 -= lr * dW2;  b2 -= lr * db2
    W1 -= lr * dW1;  b1 -= lr * db1

    if epoch % 1000 == 0:
        print(f"epoch {epoch:4d}  loss {loss:.4f}")

print("Predictions:", y_hat.round(3).ravel())   # approx [0, 1, 1, 0]

The same model in PyTorch

import torch, torch.nn as nn

model = nn.Sequential(
    nn.Linear(784, 128), nn.ReLU(), nn.Dropout(0.2),
    nn.Linear(128, 64),  nn.ReLU(),
    nn.Linear(64, 10)            # raw logits; CrossEntropyLoss applies softmax
)
loss_fn = nn.CrossEntropyLoss()
opt = torch.optim.Adam(model.parameters(), lr=1e-3)

for xb, yb in train_loader:
    opt.zero_grad()
    loss = loss_fn(model(xb), yb)
    loss.backward()              # backpropagation
    opt.step()                   # parameter update

13. Strengths and limitations

StrengthsLimitations
  • Universal function approximator
  • Simple, well-understood, easy to implement
  • Works on any fixed-size feature vector (tabular data)
  • Learns feature interactions automatically
  • Fast inference; scales well with GPUs
  • Ignores structure: no built-in notion of spatial or temporal order (CNNs and RNNs/Transformers handle this better)
  • Parameter count grows quickly with input size (e.g., raw images)
  • Needs feature scaling and tuning
  • Prone to overfitting on small datasets
  • Largely a "black box"; limited interpretability
  • On many tabular problems, gradient-boosted trees can still match or beat it

14. Applications

  • Tabular prediction: credit scoring, churn prediction, house price estimation, demand forecasting.
  • Classification: handwritten digits (MNIST), spam detection, medical diagnosis from clinical features.
  • Function approximation and surrogate modelling in engineering and physics.
  • Building blocks inside larger models: the feed-forward sub-layer in every Transformer block is an MLP; classification heads on top of CNN or language-model features are MLPs; the MLP-Mixer family replaces attention and convolution entirely with MLPs.
  • Reinforcement learning: policy and value networks for low-dimensional state spaces.

15. Summary

  • An MLP is a stack of fully connected layers: input → hidden layer(s) → output.
  • Each layer computes a = f(Wx + b); non-linear activations are essential.
  • Training minimizes a loss via backpropagation (chain rule) and gradient descent variants such as Adam.
  • Good results depend on proper initialization, input scaling, regularization and sensible hyperparameters.
  • Despite its simplicity, the MLP remains a core component of nearly every modern deep-learning architecture.

Found this useful? Share it, and leave a comment with your questions about neural networks.

Post a Comment

0 Comments