← Cellular Automata From First Principles

What Did the Neural CA Actually Learn?

Visualization gives us hypotheses.

Intervention gives us evidence.

If we suspect a hidden channel carries useful information, the next question is simple:

what happens if we remove it?

This chapter turns NCA interpretability into controlled experimentation.


Establish the behavioral baseline

Before touching the model, record its normal performance — with a defined suite, not an anecdote:

def score_maze(final, walls, start, goal):
    dist = final[0, 3].detach().cpu().numpy()
    optimal_dist = bfs_distance(walls, goal)

    if not np.isfinite(optimal_dist[start]):
        return ("unreachable", int(np.all(dist <= 0.05)))

    path = descend_path(dist, start, goal, walls)

    if path[-1] != goal:
        return ("unsolved", None)

    excess = len(path) - 1 - int(optimal_dist[start])
    return ("solved", excess)


def evaluate_suite(model, mazes, steps=128, runner=None):
    run = runner or rollout_frozen
    solved = optimal = unreachable_ok = 0
    excess, unreachable = [], 0

    for walls, start, goal in mazes:
        state = make_state(walls, start, goal)
        frozen = state.clone()
        final = run(model, state, steps, frozen)
        kind, value = score_maze(final, walls, start, goal)

        if kind == "unreachable":
            unreachable += 1
            unreachable_ok += value
        elif kind == "solved":
            solved += 1
            optimal += int(value == 0)
            excess.append(value)

    n = len(mazes) - unreachable

    return {
        "valid_path_rate": solved / max(1, n),
        "optimal_path_rate": optimal / max(1, n),
        "mean_excess": float(np.mean(excess)) if excess else float("nan"),
        "unreachable_accuracy": unreachable_ok / max(1, unreachable),
    }

(Verified: executes end to end on test mazes; an untrained stub scores zero throughout — the harness measures rather than flatters.)

Keep metrics such as:

valid-path rate
optimal-path rate
mean excess path length
unreachable-maze accuracy
stabilization time

Every intervention will be compared with this baseline. The intervention ladder for the chapter — what each manipulation breaks, and what a break would mean:

    flowchart LR
    B[baseline suite scores] --> Z[zero one channel]
    Z --> S[shuffle it spatially]
    S --> F[freeze it in time]
    F --> G[ablate groups]
    G --> C[compare all deltas to baseline]
  
InterventionBreaksA drop meansConfound to rule out
zero channelcontent + magnitudechannel mattered somehownetwork never used it anyway
spatial shuffleorganization onlyspatial pattern matteredmagnitude effects
freeze in timeongoing dynamicscontinued evolution matteredearly setup was enough
swap across mazesproblem-specificitycontent is maze-specificgeneric dynamics suffice
ablate groupsredundancyjoint necessitysingle-channel compensation

Zero one hidden channel

During every recurrent step, force one channel to zero.

@torch.no_grad()
def rollout_with_zeroed_channel(model, state, frozen, steps, channel):
    for _ in range(steps):
        state = model(state)
        state[:, :3] = frozen[:, :3]
        state[:, channel] = 0.0
    return state

Evaluate every hidden channel separately, reusing the suite with a channel-zeroing runner built from the rollout above:

def evaluate_with_intervention(model, mazes, steps=128, zero_channel=None):
    def runner(m, state, n_steps, frozen):
        for _ in range(n_steps):
            state = m(state)
            state[:, :3] = frozen[:, :3]
            state[:, zero_channel] = 0.0
        return state

    return evaluate_suite(model, mazes, steps=steps, runner=runner)
results = {}

for channel in range(5, 16):  # hidden channels of the 16-channel state
    results[channel] = evaluate_with_intervention(
        model,
        test_mazes,
        zero_channel=channel,
    )

A sharp performance drop tells us the channel is functionally important under this intervention.

It does not yet tell us exactly what it represents.


Rank channels by causal importance

For each metric compute a delta from baseline:

importance = baseline["valid_path_rate"] - result["valid_path_rate"]

Then rank channels.

You may find:

some channels appear almost unused
some have mild redundant roles
one or two are critical

That reveals whether the learned computation is distributed broadly or bottlenecked through a few internal variables.


Shuffle a channel spatially

Zeroing removes both content and magnitude.

A different intervention preserves the value distribution while destroying spatial organization.

def spatial_shuffle(x, channel):
    flat = x[:, channel].flatten(1)

    for batch in range(flat.shape[0]):
        order = torch.randperm(flat.shape[1], device=x.device)
        flat[batch] = flat[batch, order]

    x[:, channel] = flat.view_as(x[:, channel])
    return x

(Verified: value multiset preserved exactly, spatial layout changed.)

If zeroing has little effect but spatial shuffling destroys performance, the pattern of the signal matters more than its average level.


Freeze a channel in time

Another intervention asks whether a channel must keep evolving.

Capture it after a few steps:

frozen_value = state[:, channel].clone()

Then restore that value after every later update.

state[:, channel] = frozen_value

Try freezing at:

t = 0
t = 4
t = 8
t = 16
t = 32

This can reveal temporal roles.

A channel may matter only during early frontier propagation and become irrelevant after the global field is established.


Perturb local regions

Do not only intervene globally.

Erase a hidden channel inside one spatial region:

def erase_patch(x, channel, y0, y1, x0, x1):
    x[:, channel, y0:y1, x0:x1] = 0.0
    return x

Compare damage near:

start
goal
branch point
bottleneck
irrelevant open region

Now the intervention can reveal where an internal signal is needed.


Swap hidden state between mazes

This is a stronger test.

Take two mazes A and B after the same number of recurrent steps.

Copy one hidden channel from A into B:

state_b[:, channel] = state_a[:, channel]

Then continue B’s rollout.

If the result degrades badly, that channel contains problem-specific spatial information.

If behavior barely changes, the channel may carry generic dynamics or may be redundant.


Ablate groups, not only individuals

Neural representations are often redundant.

Two channels may compensate for one another.

So after individual ablations, test pairs and groups — for example, the top-ranked pair from the individual results.

But combinatorics grow rapidly.

Use the individual results to prioritize combinations instead of exhaustively trying every subset.


Test whether the visible output is doing hidden work

Sometimes a channel we think is merely an output participates in recurrence.

For example, if the distance prediction channel is fed back into perception at the next step, it is part of the computation.

Zero it after each step and evaluate again.

This separates:

readout-only representation

from:

recurrent working state

That distinction matters whenever outputs are embedded inside the automaton state.


Compare learned propagation with hand-coded BFS

Now we can ask a deeper question.

Does the internal dynamics resemble a classical wavefront algorithm?

Evidence might include:

activation arrival time tracks graph distance
critical channels move outward from the goal
freezing early propagation breaks distant cells first
spatial shuffling destroys pathfinding
larger mazes work with more recurrent steps

Together, those observations support an interpretation of iterative local propagation.

But do not overstate it.

The learned rule may implement a hybrid strategy unlike literal BFS.

solves examples
  ≠ discovered algorithm

iterative local behavior
  ≠ symbolic program

generalizes somewhat
  ≠ exact algorithmic equivalence

Build an evidence table

For each interpretability claim, record its support.

claim:
"channel 7 carries goal-distance information"

observations:
- strong correlation with BFS distance
- linear probe predicts distance accurately
- zero ablation reduces valid-path rate
- spatial shuffle damages performance
- activation propagates outward from goal

This is much stronger than naming a heatmap by eye.


What the NCA taught us

Across Part V, the NCA has forced us to revisit nearly every theme in the book:

locality
state
neighborhoods
repeated updates
emergence
computation
measurement
robustness
search
generalization

The difference is that the update rule is now learned.

Yet learning did not remove the need to understand the system.

It made careful experiments more important.

Note what Part V has and has not delivered for the maze task. It supplies the protocol: behavior on the training distribution, held-out families, out-of-distribution shifts, then correlation, probing and intervention on hidden state, each rung stronger than the last. It does not establish that any trained model learned breadth-first search. That claim would need the whole ladder, and even then the honest conclusion might be a hybrid local strategy rather than a literal BFS.


We now have the full conceptual spine

We started with a binary row and an explicit rule table.

We ended with a learned recurrent local system whose hidden states can be probed and perturbed.

The remaining work is no longer about introducing a fundamentally new kind of cellular automaton.

It is about turning everything we built into a reusable, efficient and reproducible laboratory.

Part VI therefore turns to engineering:

profiling
vectorization
GPU execution
FFT convolution
a reusable engine
reproducibility
parameter sweeps
figure generation
a laboratory that ties them together

Its governing rule is the one this part has just practiced on hidden state: do not trust a result until the implementation behind it has been checked. The next chapter starts where any optimization should: by profiling.


Research

  • Mordvintsev, A., Randazzo, E., Niklasson, E. & Levin, M. — Growing Neural Cellular Automata (Distill, 2020). The substrate every intervention here operates on: shared local rule, hidden working state, pool-trained persistence. Ablation asks what the trained system uses; this paper specifies what was trained. https://doi.org/10.23915/distill.00023

  • Earle, S., Yildiz, O., Togelius, J. & Hegde, C. — Pathfinding Neural Cellular Automata (2023). The BFS reference frame for the comparison section: classical wavefront procedures with known ground truth, against which learned propagation is tested rather than assumed. https://arxiv.org/abs/2301.06820