Deep Learning Basics

Four neurons are enough to expose the wiring

3 min read

A single neuron makes the arithmetic approachable. Four neurons make the wiring visible.

All four receive the same three input values, but each has its own weights and bias. They therefore produce four different responses to one input. That is enough to show what a layer is: several learned tests operating in parallel, not one large opinion.

read the layer row by row

Each neuron computes its own weighted sum:

weighted sum = input_1 * weight_1
             + input_2 * weight_2
             + input_3 * weight_3
             + bias

The input is shared. The parameters are not.

One row may be sensitive to the first input. Another may give more weight to the last two. A third may need its bias to move the sum above zero. The rows see the same data but apply different rules to it.

That distinction is easier to retain when the numbers remain small enough to inspect.

look at Z before looking at A

Z contains the raw weighted sums before the activation function changes them. It is the cleanest place to debug a surprising result.

If a Z value is unexpected, look at the input, that row’s weights, and its bias. ReLU has not hidden anything yet.

This also explains why two neurons can both produce zero after ReLU while representing different situations. A raw value of -1 and a raw value of -30 are both clipped to zero, but the first is much closer to becoming active.

The post-activation display is useful. The pre-activation values tell you how the network got there.

ReLU acts like a gate

The A values show what survives:

if z > 0, keep z
if z <= 0, output 0

For the current input, a positive row continues into the next layer and a non-positive row contributes zero. The weights have not changed; the signal has been filtered according to the row’s score.

That gate is the first point where the layer stops looking like four independent arithmetic exercises. It is now producing a selective representation of the input.

change one wire and predict the result

The interactive module is most useful when you make a small change rather than moving everything at once.

Change one weight in one row. The corresponding Z value should change in proportion to the matching input. If that input is zero, changing the weight should have no effect for this example. If the input is large, the same weight change should be easier to see.

That zero-input case is a useful warning. A parameter can exist in the model and still be invisible for a particular example. One observation does not exercise every connection.

Four rows give you the comparison needed to notice that.

a layer is a set of partial signals

With four neurons, the differences become concrete:

  • one row may respond strongly to the first input
  • another may depend mostly on the second and third
  • one may stay below zero because of its bias
  • another may remain active across a wider range

None of these rows understands the whole problem. Each produces a partial signal, and later layers can combine those signals into a more useful representation.

That is the structure to carry forward. A larger layer is more rows. A batch is more columns. Another layer repeats the same kind of calculation with a new set of inputs.

Use the example by following the numbers and predicting before you run it. A wrong prediction identifies the connection you have not understood yet.

Jeremy London

About Jeremy London

Engineering leader and builder in Denver. I write about AI platforms, agents, security, reliability, homelab infrastructure, and the parts of engineering work that have to survive production.