Evaluating Production LLMs Without Ground Truth
In standard software engineering, regression testing is deterministic: given input $X$, assert that output equals $Y$. But in generative systems, a model upgrade might rephrase an explanation, simplify prose, or adopt a more concise style—all while remaining fundamentally correct. Conversely, it might introduce subtle factual inversions that evade superficial semantic similarity metrics.
How do you evaluate an agent in production when a definitive, static "ground truth" answer does not exist?
The Triangulated Evaluation Framework
We designed a three-pronged continuous evaluation harness that runs asynchronously on a 5% shadow traffic stream:
┌──────────────────────┐
│ Production Request │
└──────────┬───────────┘
│
┌────────────────┴────────────────┐
▼ ▼
[Production Engine] [Adversarial Perturber]
│ │
▼ ▼
[Candidate Output] [Counterfactual Variant]
│ │
└────────────────┬────────────────┘
│
▼
[Multi-Judge LLM Jury]
│
┌───────────────────────┼───────────────────────┐
▼ ▼ ▼
[Consistency] [Factual Grounding] [Invariance Score]
1. Adversarial Counterfactual Perturbation
To measure model robustness, the evaluator takes real user queries and generates semantic counterfactuals:
- Changing numerical magnitudes.
- Inverting sentiment adjectives.
- Injecting contradictory premises.
If an LLM produces the same answer when the core premises are inverted, it reveals sycophancy or unmoored generation.
2. Multi-Perspective LLM Jury with Calibrated Rubrics
Rather than relying on a single "LLM-as-a-Judge" (which exhibits strong self-preference bias and positional bias), we deploy a jury of three orthogonal models:
- One focused strictly on factual consistency.
- One focused on reasoning chain completeness.
- One acting as an adversarial prosecutor looking for edge-case failures.
Each judge outputs a structured JSON evaluation with specific negative evidence spans before assigning a numerical score.
Automated CI/CD Regression Guard
Every pull request to our prompt registry or agent orchestration code triggers an automated evaluation run across 2,500 synthetic scenarios. If the cumulative grounding score or safety invariant drops by more than 0.5%, the deployment is automatically blocked.