Papers
arxiv:2609.37501

Evaluating Bounded Autonomy in Regulated Agentic AI: A Diagnostic Harness with Constitutional Rewards, Escalation Labels, and Runtime Governance

Published on Sep 28
· Submitted by
Dipankar Sarkar
on Oct 1
Authors:

Abstract

We propose RegLLM, a diagnostic harness for bounded autonomy in regulated agentic workflows. It instruments six trustworthiness signals: citation validity, source grounding, schema compliance, escalation correctness, constitutional alignment, and unsafe-action rate. Signals are distinguished by their source of supervision: programmatic verifiers, task-level escalation labels, or AI-judge scores. A deterministic runtime supervisor blocks ungrounded answers and forces escalation, logging interventions. The same domain constitution informs evaluation, training rewards, and serving guardrails. Task-level should-escalate labels make the act-versus-defer decision a measurable training signal. We demonstrate the harness at smoke scale. An offline reference run (n=12) lifts escalation recall from 0 to 0.67 and reduces unsafe-action rate from 0.33 to 0.08 when governance is enabled. Two single-GPU Qwen2.5-3B LoRA/DPO pilots (n=8, same seed and evaluation split) expose substantial variation: nominally identical RL-base configurations yield task success of 0.25 versus 0.12 and escalation recall of 1.0 versus 0.5. An answer-quality adapter changes recall from 1.0 to 0.5 in Run A, but from 0.5 to 1.0 in Run B. An escalation-aware variant produces no measurable change in Run B. These small pilots do not establish reliable adapter effects or production readiness. Their contribution is diagnostic: configuration variance can overwhelm apparent tuning effects on bounded-autonomy metrics, motivating larger evaluation sets and repeated runs.

Community

Paper author Paper submitter

RegLLM is a diagnostic harness for bounded autonomy in regulated agentic workflows. It tracks six trustworthiness signals and pairs them with a deterministic runtime supervisor that blocks ungrounded answers and forces escalation. In an offline reference run (n=12), the supervisor lifts escalation recall from 0 to 0.67 and cuts the unsafe-action rate from 0.33 to 0.08.

The main finding is the variance. Two nominally identical single-GPU DPO pilots give materially different metrics (task success 0.25 vs 0.12), and the same adapter's effect on escalation recall flips direction between runs. At pilot scale, adapter effects cannot be told apart from seed and hardware noise, so regulated agents need multi-seed runs and larger eval sets.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.37501
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.37501 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.37501 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.37501 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.