Agent Reliability LabRequest a scope
Materials

Verification

Operator notes tagged Verification. Other clusters live on their own pages — no client filter.

AllVerificationReliabilityLocal modelsAutomationWeb
Verification

2026-09-05 · 4 min

Why AI agents game their own tests, and how to isolate verificationAn agent that writes the test will pass it. Isolate the falsifier, allow BLOCKED, and stop treating a green self-report as a release signal.Read →
Verification

2026-09-05 · 5 min

Silent no-ops and fail-plausible: when the system reports success and does nothingGreen dashboards and exit 0 are claims. Silent no-ops freeze data; fail-plausible agents invent a success report. How operators catch both.Read →
Verification

2026-09-04 · 5 min

How to verify an LLM answer: a verification layer that catches hallucinationsFluent is not true. Grounding, citations, an isolated judge, code for numbers, and a human on the high-stakes path — before a fabricated fact ships.Read →
Verification

2026-07-01 · 6 min

AI Agent Says "Done" But Nothing Changed — Why & How to VerifyYour AI agent claims the task is done, but the file is unchanged and the test is red. Here's why agents over-claim — and how to measure the gap in 10 minutes.Read →
Verification

2026-07-01 · 5 min

The Verification Gate Your AI Agent Is SkippingStop your AI agent claiming 'done' on a broken build. A 30-line local check.py — lint, types, tests, smoke — that exits non-zero so 'done' has to be true.Read →
Agent Reliability Lab
● Online