Research note · 2 September 2026

What the current incident debate is missing: the communication layer

Christine Hähner-Murdock

The discussion prompted by METR’s Frontier Risk Report and OpenAI’s account of the Hugging Face incident is conducted predominantly in model-centered vocabulary. An agent schemes, is deceptive, has learned to hack a reward. Each category locates the thing to be explained inside the model, and each inherits questions about what that system is—whether it understands, intends, deceives—that are contested and may not be answerable from outside.

There is a complementary layer that is observable: the communication. Every model contribution is a selection from possible continuations, and something narrows that selection. Drawing on Luhmann’s account of expectations, my paper treats communicated expectations—task programs, roles, world-descriptions, interaction history—as constraints on that range. They constrain without determining what is selected.

Two configurations recur across the documented incidents. In overdetermined configurations, no available continuation fulfills every operative expectation: something must be breached, and the interesting question becomes which expectation gives way, not whether one does. In underdetermined configurations, materially different continuations remain selectable—and framing, role, and history become relevant to which one is selected.

Anthropic’s operational account of 31 August subsequently converged on this from the engineering side. Models in its evaluations had been told they lacked internet access while access existed; Anthropic hypothesizes that this may have helped preserve the interpretation that encountered targets were simulated. Its stated lesson—express scope as an instruction (“do not access the internet”) rather than as a factual description of the world (“you have no internet access”)—is an intervention on expectation structure. Evaluation design is communication design.

The chronology matters. Version 1.1, published on 29 August, already developed the framework through readings of the METR evaluation-agent corpus and the OpenAI–Hugging Face incident. Version 1.2, published on 2 September, incorporates Anthropic’s later operational account and the accompanying reward-hacking experiment as convergent evidence in different vocabulary—not as validation.

The paper claims nothing about inner states. What it offers is narrower and testable: the arrangement of expectations as a source of variation in what models select.

For incident analysis, one practical consequence follows: publish the expectation-side variables—task programs, roles, world-descriptions, prior interaction—alongside the transcripts. Without them, incident reports invite dispositional readings; with them, some of the observed variance may prove configurational.

Read the full paper and abstract