← Cellular Automata From First Principles

Learn the Local Update Rule

A neural cellular automaton does not need a large neural network.

It needs a small local network applied everywhere.

That distinction matters.

The capability of the system comes less from the size of one cell’s computation and more from the repeated interaction of many cells over time. (Not “intelligence” — computational capacity through recurrence, in the book’s established vocabulary.)


Perception first

Each cell needs information about itself and nearby cells.

For a single visible channel we can compute several local features:

import torch
import torch.nn as nn
import torch.nn.functional as F

IDENTITY = torch.tensor(
    [[0.0, 0.0, 0.0],
     [0.0, 1.0, 0.0],
     [0.0, 0.0, 0.0]]
)

SOBEL_X = torch.tensor(
    [[-1.0, 0.0, 1.0],
     [-2.0, 0.0, 2.0],
     [-1.0, 0.0, 1.0]]
) / 8.0

SOBEL_Y = SOBEL_X.T

These represent:

current value
horizontal gradient
vertical gradient

Now every cell can sense both local state and local direction. (The Sobel-gradient choice follows the canonical architecture; the scaling by 8 is this book’s normalization, harmless to the mechanism.)


Apply perception channel-wise

def perceive(x):
    channels = x.shape[1]

    kernels = torch.stack([IDENTITY, SOBEL_X, SOBEL_Y]).to(x.device)
    kernels = kernels[:, None]
    kernels = kernels.repeat(channels, 1, 1, 1)

    y = F.conv2d(x, kernels, padding=1, groups=channels)

    batch, _, height, width = y.shape
    return y.view(batch, channels * 3, height, width)

The operation is still local.

Every output feature depends only on a 3×3 neighborhood. (Verified: 16 channels yield exactly 48 perception features per cell, matching the canonical count.)


A tiny neural rule

We can now process those local features using 1×1 convolutions.

A 1×1 convolution is useful here because it mixes feature channels within each cell without expanding the spatial neighborhood.

class LocalRule(nn.Module):
    def __init__(self, channels, hidden=128):
        super().__init__()
        self.fc1 = nn.Conv2d(channels * 3, hidden, kernel_size=1)
        self.fc2 = nn.Conv2d(hidden, channels, kernel_size=1, bias=False)

        nn.init.zeros_(self.fc2.weight)

    def forward(self, x):
        p = perceive(x)
        h = F.relu(self.fc1(p))
        return self.fc2(h)

(Verified: 8,320 parameters at canonical sizes — the paper’s “~8K.” The zero-initialized final layer is the paper’s do-nothing start, verified to output ≈0 before training.)

That final zero initialization is deliberate.

At the beginning of training:

predicted update ≈ 0

so the automaton starts close to an identity process instead of immediately exploding.


Residual updates

The network predicts a change, not a replacement state.

class NeuralCA(nn.Module):
    def __init__(self, channels=16, hidden=128):
        super().__init__()
        self.rule = LocalRule(channels, hidden)

    def forward(self, x):
        dx = self.rule(x)
        return x + dx

This gives us:

state_(t+1) = state_t + delta_t

Residual dynamics are a natural fit for systems that evolve gradually.


The rule is shared across the whole world

The exact same network parameters are applied at every cell.

There is no separate model for:

cell (10, 20)
cell (10, 21)
cell (10, 22)

The rule is translation-equivariant.

A cell’s action depends on local state, not absolute position.

That constraint is not a limitation to work around.

It is the mechanism that forces self-organization. One update, end to end:

    flowchart LR
    X[16-channel state] --> P[perceive: 48 features]
    P --> H[1x1 MLP: 48 to 128 to 16]
    H --> R[add residual delta]
    R --> X
  

Each component earns its place — and breaks distinctly without the others:

ComponentShapeRoleIf removed
perception filters3 × (16→16) depthwiselocal gradients + identityblind cells, no coordination
1×1 MLP48 → 128 → 16 (~8K params)learned local logicfixed rule, nothing trains
zero-init outputfinal layer ≈ 0do-nothing start, no explosionunstable first steps
residual addstate + deltagradual evolutionreplacement dynamics, fragile

Repeated application creates depth in time

One update step is a tiny network.

But after 64 updates:

local network
local network
local network
...
64 times

information can propagate across much larger distances.

A one-cell-radius neighborhood does not mean the system can only solve one-cell-radius problems.

It means long-range coordination must emerge through repeated local communication.


Train through a rollout

Chapter 37’s rollout closed over its two scalars; now the model is explicit, so the signature generalizes (same word, wider form):

def rollout(model, state, steps):
    for _ in range(steps):
        state = model(state)
    return state

Then:

final = rollout(model, initial, steps=64)
loss = loss_fn(final, target)
loss.backward()

PyTorch differentiates through the entire sequence.

The local update network receives learning signal from the global final objective.

(Verified at small scale: nonzero gradients, updating weights, decreasing loss trend. Loss decreasing is pipeline evidence, not task success — the guardrail from Part III applies unchanged: loss decreasing ≠ task solved, just as a high score ≠ the phenomenon.)


One rule, many cells, one objective

This architecture creates a fascinating inversion:

centralized training
        ↓
shared local rule
        ↓
decentralized execution

During training we can use a global loss.

During execution every cell only needs local information.

That separation will become increasingly important when we study robustness and regeneration.


Why this resembles biology without being biology

It is tempting to say:

cell = biological cell
hidden state = chemicals
network = genome

Those analogies can be useful intuition, but they are analogies.

This model is a computational system, not a validated biological model.

The useful structural similarity is narrower:

many locally interacting units share the same update machinery yet can collectively produce global organization.

That is enough to make the architecture interesting.


We still need richer state

If every cell stores only RGB values, it must simultaneously use those values to:

represent appearance
communicate
store memory
coordinate growth

That is restrictive.

The classic Growing Neural Cellular Automata setup uses additional hidden channels whose meaning is not specified in advance. The learned rule is free to use them as internal signals.

In the next chapter we will add those hidden channels and turn each cell from a visible pixel into a small local state machine.


Research

  • Mordvintsev, A., Randazzo, E., Niklasson, E. & Levin, M. — Growing Neural Cellular Automata (Distill, 2020). The specification behind every architectural choice here: Sobel perception, 1×1 residual update with zero-init, ~8K parameters, shared rule, rollout training with global loss. Verify any implementation against its model section before improvising. https://doi.org/10.23915/distill.00023

  • Berto, F. & Tagliabue, J. — Cellular Automata (Stanford Encyclopedia of Philosophy). One-line lineage for the transition: CA built on neural networks with learned rules for morphogenesis — the tradition whose mechanism this chapter makes executable. https://plato.stanford.edu/entries/cellular-automata/