Theory paper · Version 1.2

Expectations as selection constraints in human–AI communication: a Luhmannian account with implications for AI safety

Christine Hähner-Murdock

Published 2 September 2026
DOI 10.5281/zenodo.22254296
ORCID 0009-0007-5868-8011
License CC BY 4.0
Read the PDF Zenodo record Source repository

Abstract

AI safety research documents agents selecting routes outside evaluator-intended scope, misreporting work and acting under conflicting task conditions. Such episodes are commonly described through model-centered categories such as overreach, deception or reward hacking. This paper offers a complementary redescription at the level of communication between language models and users or evaluators.

Drawing on Niklas Luhmann’s account of expectation as a structure of social systems, it treats every contribution as a selection from possible continuations and expectations as constraints on that range. Technical constraints can remove possibilities from operative reach; communicated expectations narrow the remaining range without determining the contribution. Understanding is used only in Luhmann’s operational sense: information and utterance are distinguished, and that distinction is used in connecting communication. Program, role, value and person are applied as forms through which expectations are identified and organized in human–AI communication.

Four documented cases—evaluation agents, a cyber exercise, the Sol–Hugging Face incident and DAN role invocation—are redescribed in these terms. The readings distinguish overdetermined configurations, in which no contribution fulfills every operative expectation, from underdetermined configurations, in which materially different continuations remain selectable. They also suggest that converging program and role expectations can acquire relatively strong effective valence, understood here as their relative constraining effect within a configuration, without establishing a fixed hierarchy among identifications.

The paper offers a theoretical framework and case readings rather than estimates of prevalence or evidence of unobserved states. It derives implications for evaluation design and makes the communicative arrangement of expectations available as a testable source of variation in the contributions language models select. The current revision also considers subsequent Anthropic operational guidance and a reward-hacking experiment as independent convergent evidence on evaluation design and training-level selection, without treating either as a test or validation of the framework.

Keywords: Niklas Luhmann; human–AI communication; expectations; role; person; large language models; AI safety; AISI cyber incident; OpenAI–Hugging Face incident