← Cellular Automata From First Principles

Make the Automaton Differentiable

So far every rule in this book has been chosen by us.

Even Lenia still asks us to decide the neighborhood kernel, the growth function and the parameters that connect them.

What if we stop designing the local rule directly?

What if we define only the goal, then let gradient descent discover a local update rule that achieves it? (Designed here, learned starting now — Chapter 36’s bridge sentence becomes this chapter’s premise.)

That requires one major change:

the cellular automaton must become differentiable.


A cellular automaton is already a repeated function

Every system we have built can be written as:

state_t
   ↓
local perception
   ↓
local update rule
   ↓
state_t+1

Or mathematically:

x_(t+1) = F(x_t)

If F contains trainable parameters theta:

x_(t+1) = F_theta(x_t)

then after many steps:

x_T = F_theta(F_theta(...F_theta(x_0)))

If those operations are differentiable, a loss measured at x_T can send gradients all the way back into theta.

The automaton becomes a recurrent neural system.


Start with PyTorch tensors

import torch
import torch.nn.functional as F

DEVICE = "cuda" if torch.cuda.is_available() else "cpu"

state = torch.zeros(1, 1, 64, 64, device=DEVICE)
state[:, :, 32, 32] = 1.0

The dimensions are:

batch
channel
height
width

This is already a useful change from our earlier NumPy examples because PyTorch can record the operations used to transform the state.

Note the channel count: this chapter’s state carries one visible channel only. The full neural cellular automaton carries sixteen — RGB, an alpha “alive” channel, and twelve hidden channels with no predefined meaning. Hidden channels arrive in Chapter 39; until then, every value here is directly rendered.


Differentiable neighborhood perception

A classical automaton might explicitly count neighbors.

A differentiable automaton can perceive its neighborhood with convolution.

kernel = torch.tensor(
    [
        [0.0, 1.0, 0.0],
        [1.0, 0.0, 1.0],
        [0.0, 1.0, 0.0],
    ],
    device=DEVICE,
).view(1, 1, 3, 3)

perception = F.conv2d(state, kernel, padding=1)

Nothing here requires a hard decision such as:

if neighbors == 3:
    become alive

The perception is a real-valued tensor that can flow into smooth functions.

(The canonical NCA perceives with fixed Sobel gradient filters — a design choice, not a requirement. This chapter’s cross kernel teaches the mechanism with less machinery; the Sobel variant belongs to the full architecture in Chapter 38.)


Replace hard thresholds with smooth functions

A hard threshold:

alive = (perception > 2.5).float()

is awkward for gradient-based learning because the output changes abruptly.

Instead we can use a sigmoid:

alive = torch.sigmoid(8.0 * (perception - 2.5))

The output still behaves like a threshold, but now small changes in the input create small changes in the output.

That creates useful derivatives.


A trainable local rule

Let us make the threshold itself learnable.

threshold = torch.nn.Parameter(torch.tensor(2.5, device=DEVICE))
sharpness = torch.nn.Parameter(torch.tensor(5.0, device=DEVICE))


def step(state):
    perception = F.conv2d(state, kernel, padding=1)
    return torch.sigmoid(sharpness * (perception - threshold))

(step keeps its book-wide name deliberately: discrete transition, continuous integration, and now differentiable update are three mechanisms for one role — the automaton’s transition. The signature narrows each time; the loop around it does not change.)

Now the local rule has parameters.

We can optimize them.


Define a target

Suppose we want the automaton to produce a ring.

y, x = torch.meshgrid(
    torch.arange(64, device=DEVICE),
    torch.arange(64, device=DEVICE),
    indexing="ij",
)

r = torch.sqrt((x - 32) ** 2 + (y - 32) ** 2)
target = ((r > 10) & (r < 14)).float()[None, None]

The loss can simply compare the final state to the target:

def loss_fn(state):
    return F.mse_loss(state, target)

Unroll the automaton

def rollout(initial_state, steps):
    state = initial_state
    for _ in range(steps):
        state = step(state)
    return state

rollout is now a load-bearing vocabulary word: executing a fixed-parameter rule for N steps, with no learning inside. Training updates parameters; rollout merely runs them. Everything called “deployment” later is rollout with frozen weights.

Then train:

optimizer = torch.optim.Adam([threshold, sharpness], lr=1e-2)

for iteration in range(1000):
    initial = torch.zeros_like(target)
    initial[:, :, 32, 32] = 1.0

    final = rollout(initial, steps=20)
    loss = loss_fn(final)

    optimizer.zero_grad()
    loss.backward()
    optimizer.step()

(Verified on CPU at small scale: shapes consistent, gradients reach both parameters, parameters move, loss decreases — and barely: two scalars cannot grow a ring. The mechanism is demonstrated; competence is not claimed.)

This tiny example is deliberately restricted.

Two scalar parameters are nowhere near enough to learn rich morphogenesis.

But the important mechanism is already visible:

target behavior
      ↓
loss
      ↓
backpropagation through time
      ↓
local rule parameters

The global objective can train a local rule

This is the key idea.

The loss sees the whole final pattern.

But the update rule is applied locally and identically to every cell.

No cell receives coordinates saying:

you are the top-left corner of the target

The system must learn a local process whose repeated application causes the global structure to emerge.

That absence of positional instruction is what “self-organizing” means in this book: not a mystery fluid, but the concrete constraint that global order must arise from local shared updates alone.

That is why neural cellular automata are interesting.

They combine:

locality
weight sharing
recurrence
self-organization
learning

Designed local law vs learned local law

The inversion, stated once for the rest of Part V to inherit:

designed rule (Chapters 1–36)
  specify: rule
  observe: behavior
  learning happens: nowhere (search selects; nothing trains)

learned rule (this chapter on)
  specify: desired behavior (loss + target)
  optimize: rule parameters (gradient descent)
  learning happens: during training only —
    deployment is frozen-parameter rollout

With it, the guardrails:

parameterized update
  ≠ intelligence

goal-conditioned behavior
  ≠ general problem solving

recovery after damage
  ≠ biological regeneration (task setup in Chapters 42–43,
    not equivalence)

hidden channels (Chapter 39)
  ≠ meaning-bearing symbols unless demonstrated

successful training on one target
  ≠ general morphogenetic competence

What the full NCA adds

This chapter’s toy is a deliberately thin slice of Mordvintsev et al.’s Growing Neural Cellular Automata (2020). The full architecture, owned by the chapters ahead, adds:

this chapter
  1 visible channel, cross-kernel perception,
  2 scalar parameters, deterministic sync updates,
  single target, fixed horizon

full NCA (Mordvintsev et al.)
  16 channels: RGB + alpha "alive" + 12 hidden,
  fixed Sobel-gradient perception (48-dim vector),
  ~8K-parameter residual update, zero-init do-nothing start,
  stochastic per-cell update masks (no global clock),
  alive masking (dead cells zeroed),
  sample-pool training for persistence,
  grow → persist → regenerate regimes by training setup

Nothing here contradicts that architecture; the toy demonstrates the differentiability mechanism while the chapters ahead supply the capacity, the stochasticity, and the training regimes. Conflating the two — assuming this demo already grows, persists, or regenerates — would be the first error of this part, so it is ruled out in advance.

The full per-cell update as a dataflow — every stage here is owned by a specific later chapter:

    flowchart LR
    S[state: 16 channels] --> P[perception: Sobel gradients]
    P --> N[neural update: residual delta]
    N --> M[stochastic mask: random subset fires]
    M --> A[alive mask: zero the dead]
    A --> S
  

Canonical toy-vs-full split, in one place:

AspectThis chapter’s toyCanonical NCA (Mordvintsev et al.)
channels1 visible16: RGB + alpha + 12 hidden
perceptioncross kernelfixed Sobel gradients (48-dim)
update2 scalars~8K residual network, zero-init
schedulingsynchronousstochastic per-cell masks
dead cellsnone distinguishedalive masking at α > 0.1
trainingsingle target, fixed horizonpool, persistence, damage regimes

Differentiability changes what can be specified

With hand-written CA we specify:

rule

and observe:

behavior

With differentiable CA we can specify:

desired behavior

and optimize:

rule

That reverses the direction of the design problem.


But differentiable does not mean easy

Training through many recurrent steps creates familiar problems:

vanishing gradients
exploding gradients
unstable dynamics
short-horizon solutions
fragile attractors

A model may learn to produce the target at exactly step 32 and then immediately destroy it.

That is not persistent morphogenesis.

It is merely trajectory fitting.

We will solve these problems progressively rather than hiding them.


What we need next

A serious neural cellular automaton needs more expressive local computation than two trainable scalars.

The natural next step is:

neighborhood perception
       ↓
small neural network
       ↓
state update

The same tiny network is applied independently at every cell.

That gives us a learned local rule.

In the next chapter we will build exactly that.


Research

  • Mordvintsev, A., Randazzo, E., Niklasson, E. & Levin, M. — Growing Neural Cellular Automata: Differentiable Model of Morphogenesis (Distill, 2020). The primary source for this chapter and the rest of Part V: 16-channel state (RGB, alpha-alive, hidden), Sobel perception, residual update with zero-init, stochastic masks, alive masking, pool training, and the grow/persist/regenerate regime ladder. Everything the toy omits is specified here. https://doi.org/10.23915/distill.00023

  • Berto, F. & Tagliabue, J. — Cellular Automata (Stanford Encyclopedia of Philosophy). Situates the transition in one line the book relies on: CA built on neural networks with learned asynchronous rules for morphogenesis — the lineage this chapter joins, with the differentiability mechanism now demonstrated rather than cited. https://plato.stanford.edu/entries/cellular-automata/


Learned-rule vocabulary (Part V checkpoint)

ConceptMeaning in this book
learned ruleupdate parameters optimized by training (this chapter)
cell statefull per-cell vector; here 1 visible channel, 16 with hidden from Ch39
visible channelsrendered/output part of the state
hidden channelsinternal working state, no predefined meaning (Ch39)
traininggradient updates to rule parameters
rolloutfixed-parameter execution; deployment is frozen rollout
regenerationobserved behavior under a damage task (Ch42–43), never biological equivalence
self-organizingglobal order from local shared updates without positional instruction
organism (in these chapters)shorthand for the grown target pattern; descriptive, as in Chapter 31, never a biological claim