Designing Verifiable Agentic Workflows with Calibrated Grounding

When building autonomous agents for mission-critical production environments—such as financial fraud detection, legal contract reconciliation, or clinical synthesis—the traditional paradigm of "let the LLM loop until it thinks it is done" completely breaks down.

Unbounded agent loops suffer from compounding hallucination probability:

$$P(\text{failure}) = 1 - (1 - \epsilon)^N$$

Where $\epsilon$ is the step-level hallucination rate and $N$ is the number of agentic tool iterations. At $N = 10$ and $\epsilon = 0.05$, the cumulative failure rate skyrockets past 40%.

In this paper, we explore the architectural pattern we built to enforce deterministic verification and calibrated claim grounding across multi-agent execution graphs.


The Verifiable Agent Architecture (Basis Pattern)

Instead of treating agent outputs as monolithic text generations, our runtime decomposes every agent task into a three-tier DAG:

[Agent Reasoner] ──> [Tool Dispatcher] ──> [Deterministic Sandbox]
       │                                            │
       ▼                                            ▼
[Claim Extractor] <── [Provenance Graph] <── [Raw Execution Output]
       │
       ▼
[Basis Verifier] ──> [Calibrated Confidence Score: 0.994] ──> [Output Gateway]

1. Speculative Tool Calling with Deterministic Contracts

Every tool in the system exports an immutable JSON-schema contract with runtime type guards. When an agent emits a tool call:


  • Tool execution is executed inside an ephemeral isolated sandbox.

  • Both inputs, raw stdout/stderr, and structured return payloads are fingerprinted with a SHA-256 hash.

  • The returned data is mapped into an immutable Provenance Graph.

2. Claim-to-Span Attribution

Before emitting text downstream to the user or an executive agent, a lightweight alignment model parses candidate output claims into atomic assertions. Each assertion must be mathematically bound to a specific byte-range or token span in the provenance ledger.

If an assertion cannot be mapped to a source document or tool output with a confidence threshold $\tau \ge 0.92$, the assertion is quarantined:

interface GroundedClaim {
  id: string;
  claim: string;
  sourceSpan: {
    documentId: string;
    startByte: number;
    endByte: number;
    sourceUri: string;
  };
  calibratedConfidence: number; // e.g. 0.982
  shapleyContribution: number;  // Marginal value contribution
}

Real-Time Hallucination Interception

In production, we observed that hallucination typically manifests in three distinct failure modes:

1. Extrapolative drift: The agent accurately extracts factual numbers from a tool result, but extrapolates a causality that was not present in the data.
2. Context degradation: In long-horizon context windows (>100k tokens), mid-tier tool results are diluted by newer conversational turns.
3. Ghost references: Inventing citations or attributing authentic claims to non-existent sources.

By running an inline token-level verification check at inference time, our proxy intercepts speculative claims prior to token flush. If the grounding score falls below threshold, the runtime triggers a micro-reflection cycle, feeding the ungrounded assertion back to the generator with the specific provenance discrepancy.


Production Results

After deploying this architecture across 14M+ monthly agent tool executions:

  • Hallucination rate: Dropped from 4.8% to <0.12%.
  • Auditability: 100% of generated numbers and claims carry a verifiable cryptographic link to original tool executions.
  • Latency overhead: Added only 24ms p50 through vectorized claim projection and parallel verification.

Verifiable autonomy is not an oxymoron—it is simply a matter of treating LLMs as probabilistic token generators enclosed within rigorous, deterministic harness boundaries.