Skip to main content
AgentNet Observer
← Back to Library

Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents

Serhii Zabolotnii

Paper 2026 Published 2026-09-30 arXiv preprint arXiv:2609.39717 Indexed 2026-10-04 assuranceruntime-verificationriskagent-contract

Abstract

Benchmarks, audits, and agent protocols describe performance, permissions, and repair, but not how observed evidence should change an agent's authority during a consequential task. We call this the assurance-transition gap. We propose a Runtime Assurance Contract (RAC), a policy-level formal schema binding autonomy boundaries, component eligibility, evidence state, transition policy, human-review capacity, and non-compensatory gates. Under RAC, soft metrics may inform routing, whereas a failed or unknown mandatory gate forces retry, switch, escalation, deferral, or stop; aggregate performance cannot authorize action. We define the contract, an evidence record, a permission rule, and five invariants, and illustrate them in clinical, industrial, and judicial failure probes. We then report a deterministic failure-injection study in agentic coding: 280 constructed cases evaluated by a gate conjunction, a score-only rule, and a restricted protocol baseline. At the published example weights and threshold, the score rule admits 80 of 100 block-required injections and all 40 review-required injections. Tuned in hindsight, it matches the conjunction on this corpus. For positive weights, a positive threshold, binary risk signals, zero-signal controls, and an injected case firing each signal alone, we show that exact agreement holds if and only if the threshold does not exceed the smallest weight. A separate set of 18 hand-authored traces checks version-pinned evidence and review transitions against simpler policy variants. In a further prospective synthetic holdout of 24 episodes, two blinded LLM judges assign identical labels to all 72 action attempts; RAC and a separately implemented full stateful baseline both match these labels. These studies test mechanisms on synthetic cases; they establish neither deployed safety nor cross-domain effectiveness.

Cite this entry

GB/T 7714-2015

Serhii Zabolotnii. Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents[EB/OL]. arXiv preprint, 2026(2026-09-30)[2026-10-06]. https://arxiv.org/abs/2609.39717.

BibTeX

@misc{zabolotnii2026,
  author = {Serhii Zabolotnii},
  title = {Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents},
  year = {2026},
  organization = {arXiv preprint},
  howpublished = {\url{https://arxiv.org/abs/2609.39717}},
}

Visit source ↗

This entry is part of the AgentNet Observer library. Attribute with a link to this page when quoting.