Restricted Boltzmann Machine(RBN)

Boltzmann Machine(RBN) in deep learning

Restricted Boltzmann Machine (RBM): A Complete, Detailed Guide

Architecture, energy function, training with Contrastive Divergence, code, variants and applications

In one line: A Restricted Boltzmann Machine (RBM) is a two-layer, undirected, stochastic neural network that learns the probability distribution of its input data by modelling the relationship between visible units and hidden units.

Table of Contents

  1. What is a Restricted Boltzmann Machine?
  2. Why "Restricted"? Architecture
  3. The Energy Function
  4. Probability Distribution and Conditional Probabilities
  5. How an RBM Learns (Training)
  6. Contrastive Divergence (CD-k) Step by Step
  7. Python Implementation (NumPy)
  8. Variants of RBM
  9. Deep Belief Networks
  10. Applications
  11. Advantages and Limitations
  12. RBM vs Autoencoder
  13. Practical Tips and Hyperparameters
  14. Conclusion

1. What is a Restricted Boltzmann Machine?

A Restricted Boltzmann Machine (RBM) is a generative, energy-based neural network. It was originally invented by Paul Smolensky in 1986 under the name Harmonium, and became famous in the mid-2000s when Geoffrey Hinton showed that it could be trained efficiently using an algorithm called Contrastive Divergence. RBMs played a major role in the early days of deep learning, because stacking them allowed researchers to pre-train deep networks layer by layer.

The RBM is a simplified version of the general Boltzmann Machine. In a general Boltzmann Machine every unit can connect to every other unit, which makes learning extremely slow. The RBM removes many of those connections, which makes training practical.

Key properties of an RBM:

  • Unsupervised: it learns from unlabeled data.
  • Generative: once trained, it can generate new samples similar to the training data.
  • Stochastic: units turn on or off with a probability, not deterministically.
  • Undirected: connections have no direction; information flows both ways.
  • Feature learner: hidden units learn useful features (patterns) of the input.

2. Why "Restricted"? The Architecture

An RBM has exactly two layers:

  • Visible layer (v): the units that represent the observed data (for example, pixels of an image).
  • Hidden layer (h): the units that capture latent features or dependencies in the data.

The "restriction" is that there are no connections between units in the same layer. Every visible unit connects to every hidden unit (a complete bipartite graph), but visible units do not connect to each other, and hidden units do not connect to each other.

Component Symbol Meaning
Visible units vi (i = 1…m) Observed input values
Hidden units hj (j = 1…n) Learned latent features
Weight matrix W (m × n) Connection strength between vi and hj
Visible bias ai Bias of visible unit i
Hidden bias bj Bias of hidden unit j

Simple picture:

   Hidden layer     (h1)   (h2)   (h3)
                      |  \  /  |  \  /  |
                      |   \/   |   \/   |      every visible unit connects
                      |   /\   |   /\   |      to every hidden unit,
                      |  /  \  |  /  \  |      but NO same-layer links
   Visible layer    (v1)   (v2)   (v3)   (v4)

Why this restriction matters: because there are no intra-layer connections, the hidden units are conditionally independent given the visible units, and vice versa. This lets us compute and sample an entire layer in parallel, in a single step. This is what makes RBM training efficient.

3. The Energy Function

RBMs are energy-based models. Every possible configuration (v, h) of the network is assigned a scalar value called its energy. Low energy means the configuration is likely; high energy means it is unlikely. For a binary RBM:

E(v, h) = − Σi aivi − Σj bjhj − ΣiΣj vi wij hj

In matrix form: E(v, h) = − aTv − bTh − vTWh

Intuition: if a weight wij is large and positive, then having vi = 1 and hj = 1 together lowers the energy, making that combination more probable. Learning means adjusting weights and biases so that configurations resembling the training data get low energy.

4. Probability Distribution and Conditional Probabilities

Joint probability. The energy is turned into a probability using the Boltzmann (Gibbs) distribution:

P(v, h) = e−E(v,h) / Z     where   Z = Σv,h e−E(v,h)

Z is the partition function, a normalizing constant summing over all possible configurations. For realistic sizes, Z is computationally intractable (it has 2m+n terms). This single fact is the root of why training an RBM is hard.

Marginal probability of the data (what we actually want to maximize):

P(v) = (1/Z) Σh e−E(v,h)

Conditional probabilities (the useful part). Thanks to the bipartite structure, they factorize neatly:

P(hj = 1 | v) = σ( bj + Σi vi wij )

P(vi = 1 | h) = σ( ai + Σj wij hj )

where σ(x) = 1 / (1 + e−x) is the logistic sigmoid function. Each hidden unit looks at the visible layer and "fires" with a probability given by the sigmoid of its weighted input; the same applies in reverse for visible units. Because of conditional independence, all hidden units can be sampled at once, then all visible units at once.

Free energy. Summing out the hidden units gives a handy quantity, the free energy: F(v) = − aTv − Σj log(1 + ebj + (WTv)j). Then P(v) = e−F(v) / Z. Free energy is often used to monitor training and to detect anomalies (anomalous inputs get high free energy).

5. How an RBM Learns

The goal is to adjust W, a and b to maximize the log-likelihood of the training data, i.e. make the training samples have high probability (low free energy). The gradient of the log-likelihood with respect to a weight has a beautiful form:

∂ log P(v) / ∂ wij = ⟨vihj⟩data − ⟨vihj⟩model
  • Positive phase, ⟨vihj⟩data: correlation measured when the visible units are clamped to real training data. Easy to compute.
  • Negative phase, ⟨vihj⟩model: correlation measured when the network runs freely and generates its own "fantasy" samples. Hard to compute exactly because it requires sampling from the full model distribution.

Intuitively: the positive phase lowers the energy of real data, and the negative phase raises the energy of everything the model currently believes in, which pushes the model toward the data distribution.

The update rules are:

  • Δwij = ε ( ⟨vihj⟩data − ⟨vihj⟩model )
  • Δai = ε ( ⟨vi⟩data − ⟨vi⟩model )
  • Δbj = ε ( ⟨hj⟩data − ⟨hj⟩model )

where ε is the learning rate.

6. Contrastive Divergence (CD-k) Step by Step

Running a Gibbs chain until it reaches equilibrium to estimate the negative phase would take far too long. Hinton's shortcut, Contrastive Divergence, starts the chain at a training example and runs it for only k steps (often k = 1). This gives a biased but very effective approximation.

The CD-1 algorithm for one training vector v(0):

  1. Positive phase: compute P(h | v(0)) = σ(b + WTv(0)) and sample a hidden state h(0).
  2. Reconstruction: compute P(v | h(0)) = σ(a + Wh(0)) and sample a reconstructed visible vector v(1).
  3. Negative phase: compute P(h | v(1)) (using probabilities rather than samples here reduces noise).
  4. Update parameters:
    • W ← W + ε ( v(0)P(h|v(0))T − v(1)P(h|v(1))T )
    • a ← a + ε ( v(0) − v(1) )
    • b ← b + ε ( P(h|v(0)) − P(h|v(1)) )
  5. Repeat over mini-batches and many epochs.

Variants of the training algorithm:

  • CD-k: run k Gibbs steps; larger k gives a better gradient estimate but costs more.
  • Persistent CD (PCD): keep the Gibbs chain running between updates instead of restarting from data; usually gives a better model.
  • Fast PCD, Parallel Tempering: more advanced samplers that mix better.

7. Python Implementation (NumPy)

A compact binary RBM trained with CD-1:

import numpy as np

def sigmoid(x):
    return 1.0 / (1.0 + np.exp(-x))

class RBM:
    def __init__(self, n_visible, n_hidden, lr=0.1):
        self.W = np.random.normal(0, 0.01, (n_visible, n_hidden))
        self.a = np.zeros(n_visible)   # visible bias
        self.b = np.zeros(n_hidden)    # hidden bias
        self.lr = lr

    def sample_h(self, v):
        p_h = sigmoid(v @ self.W + self.b)
        return p_h, (np.random.rand(*p_h.shape) < p_h).astype(float)

    def sample_v(self, h):
        p_v = sigmoid(h @ self.W.T + self.a)
        return p_v, (np.random.rand(*p_v.shape) < p_v).astype(float)

    def train_batch(self, v0):
        # Positive phase
        p_h0, h0 = self.sample_h(v0)
        # Reconstruction
        p_v1, v1 = self.sample_v(h0)
        # Negative phase
        p_h1, _ = self.sample_h(v1)

        n = v0.shape[0]
        self.W += self.lr * (v0.T @ p_h0 - v1.T @ p_h1) / n
        self.a += self.lr * np.mean(v0 - v1, axis=0)
        self.b += self.lr * np.mean(p_h0 - p_h1, axis=0)

        return np.mean((v0 - p_v1) ** 2)   # reconstruction error

# Example usage with random binary data
data = (np.random.rand(500, 20) > 0.5).astype(float)
rbm = RBM(n_visible=20, n_hidden=10, lr=0.1)

for epoch in range(50):
    err = rbm.train_batch(data)
    if epoch % 10 == 0:
        print(f"Epoch {epoch}: reconstruction error = {err:.4f}")

Alternatively, scikit-learn provides a ready-made implementation: sklearn.neural_network.BernoulliRBM.

8. Variants of RBM

Variant Description
Bernoulli-Bernoulli RBM Both layers binary. The standard form described above.
Gaussian-Bernoulli RBM Visible units are real-valued (Gaussian), hidden units binary. Used for continuous data such as image pixels or audio features.
Replicated Softmax RBM Models word-count data in documents (topic modelling).
Conditional RBM (CRBM) Adds conditioning inputs from past time steps; used for time series like motion capture data.
Convolutional RBM Shares weights across spatial locations, suitable for images.
Discriminative RBM Includes label units so the RBM can be used directly as a classifier.
Spike-and-Slab RBM Hidden units have both a binary "spike" and a real-valued "slab" for richer image modelling.

9. Deep Belief Networks (DBN)

A Deep Belief Network is built by stacking RBMs. The procedure, called greedy layer-wise pre-training, works like this:

  1. Train the first RBM on the raw input data.
  2. Freeze it, and use its hidden-unit activations as the "data" for training a second RBM.
  3. Repeat for as many layers as desired.
  4. Optionally fine-tune the whole stack with backpropagation on a supervised task.

In 2006 this was the breakthrough that made deep networks trainable at a time when plain backpropagation struggled with vanishing gradients. Later techniques (ReLU, better initialization, batch normalization, dropout) made pre-training largely unnecessary, but DBNs remain historically important.

10. Applications

  • Collaborative filtering / recommender systems: an RBM was part of the winning approach in the Netflix Prize competition, modelling user-movie ratings.
  • Dimensionality reduction and feature extraction: hidden activations serve as compact learned features.
  • Pre-training deep networks (via DBNs).
  • Image and speech recognition: early deep speech models used DBN pre-training.
  • Anomaly detection: unusual inputs have high free energy.
  • Topic modelling of documents.
  • Generative modelling: sampling new data such as handwritten digits resembling MNIST.
  • Quantum many-body physics: RBMs are used as compact variational representations of quantum states (neural-network quantum states).

11. Advantages and Limitations

Advantages Limitations
• Learns useful features without labels
• Efficient parallel Gibbs sampling
• Can both generate and represent data
• Handles missing data naturally
• Simple, elegant probabilistic foundation
• Partition function is intractable
• Training is sensitive to hyperparameters
• CD gradient is biased
• Hard to evaluate model quality (no exact likelihood)
• Largely superseded by VAEs, GANs and diffusion models

12. RBM vs Autoencoder

Aspect RBM Autoencoder
Nature Stochastic, probabilistic Deterministic (standard version)
Connections Undirected, one shared weight matrix Directed, separate encoder and decoder
Objective Maximize likelihood (minimize energy) Minimize reconstruction error
Training Contrastive Divergence Backpropagation
Generative? Yes, by Gibbs sampling Not by default (VAEs are)

13. Practical Tips and Hyperparameters

  • Number of hidden units: start with a number comparable to or somewhat larger than the visible units for feature learning; use fewer for compression.
  • Learning rate: typically 0.001 to 0.1; too high makes training diverge.
  • Weight initialization: small random values (for example, Gaussian with standard deviation 0.01); initialize visible biases from the log-odds of the data means.
  • Mini-batches: sizes of 10 to 100 work well.
  • Momentum and weight decay: momentum speeds up learning; L2 weight decay (about 0.0001) prevents large weights.
  • Sparsity: encouraging sparse hidden activations often produces more interpretable features.
  • Monitoring: reconstruction error is a rough guide only, because it is not the true objective; also track free energy gap between training and validation data to detect overfitting, and visualize learned weights as images.
  • Use probabilities, not samples, in the final reconstruction step of CD to reduce variance.

14. Conclusion

The Restricted Boltzmann Machine is a beautiful blend of probability theory, statistical physics and neural networks. By restricting connections to a bipartite graph, it keeps the expressive power of an energy-based model while allowing fast, layer-wise parallel inference. Though newer generative models have taken the spotlight, RBMs remain an excellent way to understand energy-based learning, Gibbs sampling, and the origins of deep learning, and they are still used in recommender systems and quantum physics research.

Key takeaways:
  • RBM = visible layer + hidden layer, fully connected between layers only.
  • Defined by the energy E(v,h) = −aTv − bTh − vTWh.
  • Layers are conditionally independent, so sampling is fast.
  • Trained with Contrastive Divergence to approximate the log-likelihood gradient.
  • Stacked RBMs form Deep Belief Networks.

Tags: Restricted Boltzmann Machine, RBM, Deep Learning, Machine Learning, Generative Models, Contrastive Divergence, Deep Belief Network

Post a Comment

0 Comments