PGD-AI · IIT Jodhpur · Group 18

Deep Belief
Networks

An interactive breakdown of Hinton et al. (2006) — how greedy layer-wise pretraining unlocked deep learning and achieved 1.25% error on MNIST.

Restricted Boltzmann Machines Greedy Pretraining 1.25% MNIST Error State of the Art · 2006
▼ scroll to explore
Section 01

Network Architecture

784 → 500 → 500 → 2000 → 10  ·  3 RBM layers + softmax output

Hover over a layer for details  ·  Neurons shown proportionally

Section 02

Restricted Boltzmann Machines

The building block of DBNs — a two-layer generative model that learns representations unsupervised.

Click Forward Pass to see how RBM encodes input

🔒 The "Restricted" Part

No connections within a layer — only between visible and hidden units. This makes exact inference tractable using Gibbs sampling and Contrastive Divergence.

🔄 Contrastive Divergence

Training approximation: run one step of Gibbs sampling (not full MCMC). Positive phase captures data statistics; negative phase captures model's reconstruction. Update weights by the difference.

📊 Energy Function

E(v,h) = −bᵥᵀv − bₕᵀh − vᵀWh

The RBM assigns low energy to data configurations it has learned. Learning = making data configurations lower energy than random ones.

📐 Weight Update Rule

ΔW = α(<vhᵀ>data − <vhᵀ>recon)

Increase weights when visible and hidden units co-activate on real data; decrease when they co-activate on reconstructions.

Section 03

Two-Phase Training

The breakthrough insight: initialize with unsupervised pretraining, then fine-tune with labels.

1

Unsupervised Pretraining

Train each RBM greedily, one layer at a time. No labels needed — the network learns to reconstruct its input.

  • RBM 1: 784 ↔ 500  — learns edges, strokes from raw pixels
  • RBM 2: 500 ↔ 500  — learns curves and junctions from edges
  • RBM 3: 500 ↔ 2000 — learns digit templates from curves
  • Each layer uses the previous layer's activations as its visible data
2

Supervised Fine-Tuning

Add a 10-unit softmax output and unroll the network. Run backpropagation with digit labels to adjust all weights.

  • Weights start in a good region — pretraining found a useful basin
  • Backprop adjusts all layers jointly for classification
  • Much less prone to vanishing gradients — good initialization
  • Result: weights that both model the data and classify well

Press Animate Training to see the two phases

Section 04

MNIST Results

Hinton et al. (2006) — 60,000 training / 10,000 test images, 28×28 pixels

1.25%
DBN Test Error (best)
10.7%
Relative improvement vs SVM
17.2%
Relative improvement vs Backprop
60K
Training images used

Simulated MNIST-style digits — the network learns to classify all 10 classes

Method Test Error Bar vs DBN
Deep Belief Network (DBN) 1.25%
★ Best
SVM — Polynomial Degree 9 1.40%
+0.15%
Backprop 500-300 hidden 1.51%
+0.26%
Backprop 800 hidden 1.53%
+0.28%
Neural Net — 3 layers 2.80%
+1.55%
Neural Net — 2 layers 3.10%
+1.85%