Multilayer Perceptron: 16--16--16--16

Updated 2026-10-05 · NEXMASON ANITEX

▶ Open interactive ANITEX · equations and animations

This accessible text edition preserves the document's narrative. See the interactive edition for typeset equations, diagrams and playback.

Architecture

This separate demonstration contains 16 input neurons, two hidden layers of 16 neurons each, and 16 output neurons. Each neuron connects to every neuron in the next layer. There are 256 weighted connections per transition, 768 weights in total, and 48 bias parameters. The input layer does not have a bias or an activation function. The existing L/T/X recognition demonstration is unchanged. Here L, T and X are input examples only. The 16 output classes have no assigned meanings: these fixed illustrative weights are not a trained classifier.

Matrices and Activations

A 4-by-4 input image is flattened row by row into a 16-entry column vector. Every weight matrix has dimensions 16 by 16; every bias and activation vector has 16 entries. a^(0)=x, z^(1)=W^(1)x+b^(1), a^(1)=ReLU(z^(1)). z^(2)=W^(2)a^(1)+b^(2), a^(2)=ReLU(z^(2)). z^(3)=W^(3)a^(2)+b^(3), p_i=(z_i^(3)-m)_j=1^16(z_j^(3)-m), m=_j z_j^(3). ReLU maps negative preactivations to zero. Softmax normalizes the 16 output scores to sum to one; it does not make the model trained or its scores calibrated confidence.

Every Neuron Receives All 16 Inputs

z_i^(k)=_j=1^16W_ij^(k)a_j^(k-1)+b_i^(k). A zero activation still has all its connections, but contributes zero to the next weighted sum. Blue lines represent positive weights and orange lines represent negative weights. Moving dots show nonzero contributions during the matrix multiplication stages. Node colors and numbers show the computed values.

Reproducible Teaching Parameters

For zero-based indices i and j, layer k uses the following deterministic rule. Let r be the remainder of (7i+11j+3k) divided by 17, minus 8. The weight is r/16, except that r=0 uses 1/16. Thus every displayed connection has a nonzero weight. Hidden biases are ((i modulo 3)-1)/10; output biases are zero. These rules illustrate computation, not learned visual features.

Interactive Forward Pass

Select an example or edit individual pixels. Play, Pause, Reset, speed selection and the timeline allow inspection of each calculation stage. Expand the matrix panels to inspect all weights, biases, preactivations and activations. On narrow screens, scroll the network diagram horizontally.

Training Versus Inference

This animation performs inference only. To recognize 16 real categories, assign labels, collect representative training and evaluation data, choose a loss, and optimize the weights and biases by backpropagation. Architecture alone does not guarantee recognition accuracy.