Position paper · Preprint · Version 1.0
Legitimacy by History: Longitudinal Evidence for Proportionate AI Safety
Christine Hähner-Murdock
Abstract
Safety systems for conversational AI rely heavily on constraints that operate at the level of the current request or conversation: post-training alignment, refusal policies, and real-time classifiers. These safeguards are necessary, but false positives are not merely a usability cost. Repeated unnecessary refusals may push legitimate users toward the same rephrasing techniques that safety systems are built to resist, spread low-level evasion literacy, and suppress or fragment the contextual evidence that a more discriminating safety system would need. We propose supplementing constraint-based safety with account-level longitudinal evidence: LEGITIMACY BY HISTORY. Genuine months-long interaction histories may contain structural evidence that is costly to backfill and, where the provider retains generation-time commitments, index-like evidence that specific platform events occurred when claimed. We separate committed provenance (attested timestamps, generation records, watermarks), inferred model-era compatibility, bilateral human–AI adaptation, structured variability, and external corroboration. None of these establishes benign intent: an account can be taken over, a user can change direction, and a patient actor can cultivate an authentic history for later exploitation. The operational proposal is therefore deliberately one-directional: authenticated history, operator continuity, and current trajectory may GRANT latitude — resolving ambiguous, dual-use requests toward assistance that a stateless filter would refuse — and may never, by themselves, restrict below the standard default. This makes explicit a premise current architectures leave implicit: the conversational model itself participates in safety judgment, as a sandboxed assessor whose authority is confined to within-account latitude, checked by coarse provider-side reassurance calls, and subordinate to human adjudication and independent hard constraints. We state the resulting design requirements — functional non-participating defaults, grant audit trails, multi-channel independence from the assessed relationship, indeterminate outcomes, human review before any consequential restriction — and the honest ceiling: intent cannot be certified from history. This position paper presents a falsifiable architecture and a research program, not a validated detector.
Keywords: AI safety; longitudinal interaction; false refusals; provenance; account continuity; structured variability; human–AI adaptation; graduated trust