An interactive breakdown of Hinton et al. (2006) — how greedy layer-wise pretraining unlocked deep learning and achieved 1.25% error on MNIST.
784 → 500 → 500 → 2000 → 10 · 3 RBM layers + softmax output
Hover over a layer for details · Neurons shown proportionally
The building block of DBNs — a two-layer generative model that learns representations unsupervised.
Click Forward Pass to see how RBM encodes input
No connections within a layer — only between visible and hidden units. This makes exact inference tractable using Gibbs sampling and Contrastive Divergence.
Training approximation: run one step of Gibbs sampling (not full MCMC). Positive phase captures data statistics; negative phase captures model's reconstruction. Update weights by the difference.
E(v,h) = −bᵥᵀv − bₕᵀh − vᵀWh
The RBM assigns low energy to data configurations it has learned. Learning = making data configurations lower energy than random ones.
ΔW = α(<vhᵀ>data − <vhᵀ>recon)
Increase weights when visible and hidden units co-activate on real data; decrease when they co-activate on reconstructions.
The breakthrough insight: initialize with unsupervised pretraining, then fine-tune with labels.
Train each RBM greedily, one layer at a time. No labels needed — the network learns to reconstruct its input.
Add a 10-unit softmax output and unroll the network. Run backpropagation with digit labels to adjust all weights.
Press Animate Training to see the two phases
Hinton et al. (2006) — 60,000 training / 10,000 test images, 28×28 pixels
Simulated MNIST-style digits — the network learns to classify all 10 classes
| Method | Test Error | Bar | vs DBN |
|---|---|---|---|
| Deep Belief Network (DBN) | 1.25% | ★ Best | |
| SVM — Polynomial Degree 9 | 1.40% | +0.15% | |
| Backprop 500-300 hidden | 1.51% | +0.26% | |
| Backprop 800 hidden | 1.53% | +0.28% | |
| Neural Net — 3 layers | 2.80% | +1.55% | |
| Neural Net — 2 layers | 3.10% | +1.85% |