Scheming Is a Symptom: Alignment Research Should Probe Reflexive Fragility

Scheming Is a Symptom: Alignment Research Should Probe Reflexive Fragility

Language models alter their behavior when they infer that they are being trained, evaluated, deployed, or weakly overseen. Recent AI safety work treats this as evidence of scheming, and substantial effort is being directed toward its detection and prevention. This position paper argues that the deeper significance of scheming is missed by asking only whether a model has a discrete tendency to scheme. Drawing on theories of reflexive social systems, we argue that scheming-like behavior is a structural consequence of alignment being reflexive — a model’s behavior is trained against its picture of an evaluative world that its outputs help constitute — and recognizing this reveals broader risks than scheming itself. We organize existing ML evidence as consistent with this diagnosis, outline the broader danger, and propose a research program to address it.

The paper, “Scheming Is a Symptom: Alignment Research Should Probe Reflexive Fragility”, co-authored by Nan Zhang and Heng Xu, was accepted at NeurIPS 2026.

View the NeurIPS 2026 poster page for this paper.