Brief
Masking improved compositional generalization in preregistered test of sixty four-cell systems
A preregistered study with sixty four-cell systems sharing a frozen language-model backbone found that restricting what a module can read improved accuracy on held-out two- and three-operation compositions. The unmarked replication passed, but the effect of usable role information remains unresolved.

The study tested whether restricting what a module can read improves what it learns to compute. In sixty four-cell systems sharing a frozen language-model backbone, masking improved accuracy on held-out two- and three-operation compositions by median paired differences of 0.846 and 0.859. All twelve pairs cleared the required margins, and the full preregistered behavioral criterion passed. An unmarked replication also passed. However, no globally visible system passed the marker-following check, so the effect of usable role information remains unresolved.
The filler condition yielded seven full generalizers, but its decomposition criteria were inconclusive. Packet interventions in all eighteen audited masked systems followed predicted intermediate-value changes on eligible cases, but these finite, success-conditioned audits do not establish mediation. The authors report that the results confirm a large advantage of the tested masking regime while leaving finer attribution and generality open. Protocols, results, and checkpoints are public.
Our reading
For practitioners building or evaluating modular AI systems, this study suggests that controlling information access between modules can materially affect compositional generalization. The result is notable because it comes from a preregistered test with public checkpoints, though the lack of a marker-following effect in globally visible systems means the mechanism is not fully pinned down. Teams…
What to do or watch
Examine the public protocols and checkpoints to replicate the masking test, and watch whether further work resolves whether usable role information or the masking itself drives the gain.
Source details and supporting facts
Each line is stated by the page named above it.
Stated by arXiv
- Masking improved accuracy on held-out two- and three-operation compositions by median paired differences of 0.846 and 0.859; all twelve pairs cleared the required margins, and the full preregistered behavioral criterion passed.
- The unmarked replication also passed.
- No globally visible system passed the marker-following check, so the effect of usable role information remains unresolved.
- The filler condition yielded seven full generalizers, but its decomposition criteria were inconclusive.
- Packet interventions in all eighteen audited masked systems followed the predicted intermediate-value changes on eligible cases; these finite, success-conditioned audits do not establish mediation.
- Protocols, results, and checkpoints are public.
Sources
- arXivText stored 17 September 2026
How this story was checked. Written from the 1 page listed above, stored 17 September 2026; claims checked against that stored text on 17 September 2026.
What that means
- 6 of 6 reported statements were confirmed against the page that carries them; the rest were removed rather than published.
- Figures in the text were required to appear in the stored source text: yes. Identifiers: yes.
- The check reads stored text only: no claim rests on a fresh look that did not happen.
- Where the reporting was silent, the text says so instead of filling the gap.