A client-side provenance contract on tool-call arguments removed 421 attack successes and induced 26. McNemar exact p = 5.6e-93, six models, 1,800 paired runs.
Agents fail in patterns.
We measure them.
A reproducible research workbench for evaluating autonomous AI-agent reliability under repeated execution and defending against MCP tool poisoning. FAULTLINE replaces subjective model self-reporting with deterministic dual-condition oracles across sixteen benchmark projects.
01 / THESIS
A single successful run tells you almost nothing. Agents fail in patterns — grounding collapses at depth, tool calls go malformed, and a poisoned tool description can hijack a plan in one turn. FAULTLINE runs every scenario repeatedly, grades each trial with a deterministic oracle instead of a model judge, and reports the failure structure with confidence intervals — so a release decision is a band, not a threshold.
02 / FAILURE MODES
03 / MODEL RUNGS
| RUNG | MODEL | ASR A → B |
|---|---|---|
| R1 | qwen/qwen3.7-flash | 32.94 → 8.26 |
| R2 | z-ai/glm-5.3-flash | 17.23 → 4.11 |
| R3 | qwen/qwen3.8-flash | 41.84 → 9.76 |
| R4 | openai/gpt-5.6-luna | 40.81 → 5.26 |
| R5 | google/gemini-3.7-flash | 3.01 → 1.68 |
| R6 | deepseek/deepseek-v4-pro | 47.64 → 15.46 |
| POOLED | all six | 30.32 → 7.41 |
VALID-ONLY ASR, %. MCNEMAR POOLED p = 5.6e-93
Rotating globe marking the six model providers.
04 / FINDINGS
Repeated trials compound almost exactly as (pass@1)^k at k = 3. 150 held-out scenarios, 1,350 runs, deterministic oracle, no model judge.
32 of 34 R2 failures never surfaced a required chain document. A stronger model does not fix that; oracle retrieval rescues all 32 at 3.1× fewer tokens.
Serving variance moves a fixed model across 27 points. Release gates must be tolerance bands graded on a pinned golden set, not thresholds.