Test Generalization Beyond Training
A model can look robust while still depending on the exact conditions used during training.
So after growth, persistence and regeneration, we need a harder question:
what happens when the world changes?
Build a generalization matrix
Vary dimensions independently:
canvas size
seed position
update rate
rollout length
damage geometry
damage severity
noise level
boundary conditions
Then evaluate combinations that were not used during training. The table at the end of this chapter is executable scaffolding, not a report: every row specifies how to measure, and the verdicts get filled in by running, starting from one held-out axis at a time rather than all shifts at once.
def evaluate_condition(model, make_initial, target, steps=128):
state = make_initial()
final = rollout(model, state, steps)
return float(F.mse_loss(final[:, :4], target))
Shift the seed
If training always starts at the center, test elsewhere:
def make_seed_at(y, x, size=96, channels=16):
state = torch.zeros(1, channels, size, size, device=DEVICE)
state[:, 3:, y, x] = 1.0
return state
(Same convention as Chapter 39’s seed: alpha and hidden set, RGB dark — one seed semantics across Part V.)
Because the rule is local and shared spatially, translation should be a natural capability when boundaries do not interfere. But we should measure it, not assume it.
Change the canvas size
Train on 64×64 and evaluate on larger grids:
seed = make_seed_at(48, 48, size=96)
If the target still develops correctly, that is evidence the system has not simply encoded one fixed array position.
Inject state noise
def add_state_noise(x, sigma=0.02):
return x + sigma * torch.randn_like(x)
Test several noise levels and report performance curves rather than one anecdotal example.
Change update rates
A model trained around a 50% firing rate may fail at 20% or 90%.
rates = [0.2, 0.35, 0.5, 0.65, 0.8, 1.0]
This tests whether local coordination survives a different effective timescale.
Hold out perturbations
If training uses circular wounds, evaluate rectangles and slices.
If training uses small wounds, evaluate larger ones.
This separates:
robustness to familiar corruption
from:
robustness to new corruption
Report a table, not a victory image
For example — template only; every entry below must be measured, including the failures:
condition final loss recovery time survived
-------------------------------------------------------------
center seed ... ... ...
shifted seed ... ... ...
96x96 canvas ... ... ...
20% fire rate ... ... ...
large slice damage ... ... ...
noise sigma=0.05 ... ... ...
This is much more informative than selecting the best animation — provided “survived” means measured against the task threshold, not eyeballed.
Generalization has a boundary
A local learned rule may generalize impressively within one family of dynamics while failing abruptly outside it.
That boundary is scientifically interesting.
Do not hide it.
Map it. And keep the chapter’s central doctrine attached to every map:
success outside the training example is still not success outside the training distribution.
works on held-out examples
≠ broad generalization
works on nearby conditions
≠ out-of-distribution robustness
regeneration on unseen damage
≠ arbitrary repair
The next transition
We now have training setups and test protocols for four capabilities, each of which has to be demonstrated rather than assumed:
grow
persist
repair
survive some distribution shift
So far, generalization has meant preserving a learned form under changed conditions. There is a harder test. Instead of asking the automaton to preserve a form, we can ask local updates to carry out a computation whose answer depends on information distributed across the grid.
That is distributed computation over a grid, and the same architecture can be pointed at it.
In the next chapters we will use NCA for pathfinding and maze-like computation, then inspect what the learned system is actually doing internally.
Research
Earle, S., Yildiz, O., Togelius, J. & Hegde, C. — Pathfinding Neural Cellular Automata (2023). The maze-domain precedent for holding out whole axes: hand-coded and learned BFS/DFS NCAs, diameter computation with strong generalization, and adversarially evolved training mazes that improve out-of-distribution robustness. The standard this chapter’s matrix is measured against. https://arxiv.org/abs/2301.06820
Mordvintsev, A., Randazzo, E., Niklasson, E. & Levin, M. — Growing Neural Cellular Automata (Distill, 2020). The distribution-expansion methods this chapter generalizes: sample pools and damage-augmented training as the in-distribution widening that precedes genuine held-out testing. https://doi.org/10.23915/distill.00023