Neural Network Recognition: Matrices, Hidden Layers and Activation Functions
Updated 2026-10-05 · NEXMASON ANITEX▶ Open interactive ANITEX · equations and animations
This accessible text edition preserves the document's narrative. See the interactive edition for typeset equations, diagrams and playback.
What the Network Recognizes
This interactive example recognizes three simple 4-by-4 binary patterns: L, T and X. Click pixels to change the input, or choose a preset. Press Play to follow the arithmetic through two hidden layers and the output layer. The weights are deliberately constructed from the three templates for teaching. This is not a trained general-purpose character recognizer. A modified or unfamiliar pattern still receives scores; the largest softmax value does not prove that it belongs to a known class.
Input Matrix and Flattening
An image is represented by a matrix of pixel intensities. A value of 1 is an illuminated pixel and 0 is background. Row-major flattening produces a column vector with 16 entries. XR^44, x=vec_row(X)R^16. The architecture is 16 inputs, 6 neurons in hidden layer 1, 3 neurons in hidden layer 2, and 3 output logits.
Hidden Layer 1: Matrix Multiplication
z^(1)=W^(1)x+b^(1), W^(1)R^616. Each pair of rows scores the top and bottom halves of one template. Matching foreground pixels have positive weights, background pixels have negative weights, and pixels in the other half have zero weights. The bias centers each half-score around four matched pixels. For template pixel p and a pixel j in the corresponding half, the weight is 2p-1. The bias is the number of zero template pixels in that half minus 4. For binary inputs, the resulting preactivation equals the number of matching pixels in the half minus 4.
ReLU Activation
a^(1)=ReLU(z^(1)), ReLU(z)=(0,z). Negative scores become zero. Positive scores pass through unchanged. ReLU introduces a nonlinear operation; without nonlinear activations, several affine layers could be combined into one affine transformation.
Hidden Layer 2: Combining Features
z^(2)=W^(2)a^(1)+b^(2), a^(2)=ReLU(z^(2)). W^(2)=1&1&-0.25&-0.25&-0.25&-0.25 -0.25&-0.25&1&1&-0.25&-0.25 -0.25&-0.25&-0.25&-0.25&1&1, b^(2)=-2 -2 -2. Every Hidden 1 neuron connects to every Hidden 2 neuron: 6 times 3 equals 18 connections. Positive weights support the corresponding template; negative weights suppress competing template evidence. Each neuron sums all six weighted inputs before adding its bias. The second ReLU removes negative combined evidence.
Output Logits and Softmax
W^(3)=1.5&-0.25&-0.25 -0.25&1.5&-0.25 -0.25&-0.25&1.5, z^(3)=W^(3)a^(2). All three Hidden 2 neurons connect to all three output neurons. Output biases are zero. p_k=(z_k^(3)-m)_j(z_j^(3)-m), m=_j z_j^(3). Subtracting the largest logit improves numerical stability without changing the softmax result. The displayed prediction is the class with the largest probability. These are model scores normalized to sum to one, not independently calibrated confidence estimates.
Interactive Forward Pass
The active layer is highlighted during playback. Blue connections are positive weights; orange connections are negative weights. Moving dots show nonzero input contributions flowing into the next layer during matrix multiplication. Node intensity represents activation. The matrix panels show the actual weights, biases, preactivations and post-activation values used in the calculation.
How Learning Would Change the Weights
In a trained network, many labeled examples are used to minimize a loss such as cross-entropy. Backpropagation computes gradients of the loss with respect to weights and biases. An optimizer updates those values, and the process repeats. L=- p_y, WW-LW. This demonstration animates inference only. It does not perform gradient-based training. Real image recognition commonly uses convolutional or transformer architectures and far larger datasets.