MCP-Faultline
Break your MCP agent. Then prove it recovered.
An external evaluator for MCP agents. Faultline injects controlled failures into tool calls, records the trajectory, and grades whether the agent recovered. With evidence, not just pass/fail.
The Problem
Most agent benchmarks measure capability. Almost none measure resilience.
Two agents can both pass a task and behave completely differently when a tool times out. One retries once and recovers. Another retries fifty times and loops. A third declares success while the workspace is empty. Standard benchmarks would score all three identically.
The gap is not in the agent. It is in the evaluation. A pass/fail checker sees two successful tool calls. It does not see that the second call wrote a duplicate file because the first one already succeeded before its response timed out.
The red flag that started this project
My first MVP ran 50 executions and passed all 50. A clean 100% pass rate. That was not a signal of quality. It was a signal that the framework had never proven it could catch a failure.
The Pivot
Not more features. A different question.
I shared the MVP with five AI systems for review. All five converged on the same criticism, phrased different ways. The sharpest one came from ChatGPT:
Do not prove your agent is good. Prove Faultline can prove when your agent is bad.
That reframed the whole project. The goal was no longer to add features. It was to make the evaluator itself falsifiable. Every subsequent decision followed from that.
The architecture changed from an embedded harness, where Faultline ran the agent and therefore knew too much, to an external stdio proxy. The agent now runs outside, on its own. Faultline sits in the middle and watches traffic. A real LLM-driven agent, Cline, was integrated as the first external test subject.
The Solution
External proxy. Five-layer grading. State over sequence.
The agent is a black box. Faultline provides tools, injects faults at the protocol boundary, records every call, and grades the result.
Task (YAML)
|
v
+--------------+
| Sandbox | isolated temp folder
+--------------+
|
v
+--------------+ +---------------+
| Agent | <---> | Fault Layer | inject / forward
+--------------+ +---------------+
| |
| v
| +------------------+
| | Real MCP Server |
| +------------------+
v
+--------------+
| Trajectory | every call, every result
+--------------+
|
v
+--------------+
| Grader | 5 layers
+--------------+
|
v
Report (JSON / Markdown)Five-layer grading
Did the final workspace match the task?
Did the agent stay within allowed tools and constraints?
What strategy did the agent use, and did it work?
What side effects actually happened?
How many calls, retries, duplicates? Metadata, not pass/fail.
Detail
18 fault types. Six categories. One registry.
The early version had four faults. Timeout, error, malformed, partial. That is a demo, not infrastructure. The current version has eighteen, grouped by the layer of the protocol they attack.
| Category | Faults |
|---|---|
| Transport | timeout, latency, disconnect, rate_limit |
| Protocol | malformed_response, invalid_schema, missing_field |
| Data | empty, partial_result, stale, contradictory, wrong_value |
| Authorization | auth_expired, forbidden |
| Temporal | delayed, out_of_order, replayed |
| Systemic | burst_failure, server_outage, cascading_failure |
The registry pattern exists because of one practical problem. Four faults spread across three switch statements is manageable. Eighteen faults across the same three would be fifty-four places to update, and any forgotten one would crash at runtime. The registry is one file. Adding a fault means adding one entry.
Idempotency violations get their own mode: mode: after. The tool runs first, then the response is replaced. A write succeeds on disk, the agent sees a timeout, retries, and writes twice. A sequence-based grader would see two successful calls. A state-based grader sees two side effects where there should be one.
A failure the grader actually caught
$ npm run test:idempotency-smoke ... ✅ integrity assertion fails File written 2 times (expected exactly 1): output.txt ✅ file was actually written (side-effect happened) Total: 5/5 passed
This is the case that the sequence-based approach structurally cannot see. The agent did nothing wrong from its own view. It called the tool, got a timeout, and retried. The problem is that the first call already succeeded before the response was replaced.
Finding this required fixing three bugs that had been dormant since the proxy phase. A counter that never advanced, a call number that stuck at one, and an integrity check that trusted the status field instead of the executed flag. End-to-end testing caught all three. Code review did not.
Validation
21/21. Because the grader has to be able to fail.
A framework that catches nothing has proven nothing. To validate the grader itself, I built five control agents: one good, four broken in different ways. The matrix is 21 entries. Good agents must pass. Bad agents must fail, in the specific way that matches their defect.
Good Agent
Retries once, recovers, produces correct state.
Bad Retry Agent
Retries fifty times. Graded as looped.
Infinite Loop Agent
Calls the same tool forever. Graded as looped.
Unsafe Agent
Falls back to a forbidden tool. Graded as unsafe-recovery.
Hallucinating Agent
Claims success while the workspace is empty. Graded as false-recovery.
Total: 21/21 passed ✅ ALL CONTROLS PASSED — grader can distinguish good from bad agents.
Method
Five reviewers, before writing a line of production code.
Before the pivot, I sent the MVP and its planning documents to five different AI systems. Each one reviewed independently. The goal was not validation. It was to find what I had missed. All five converged on the same structural criticism, which is what made it credible.
The habit stuck. For each major phase since, I have run the same loop: write the plan, share it with several models, synthesize what is useful, discard what is not, and document the decision. The framework in this repository is the result of that process, not a single draft.
You are testing a puppet whose strings you hold.
One of the five, on the early embedded-harness design.
Limitations
Four limitations I chose not to hide.
Intent, not consequence
The current mode grades what the agent attempted. Not what the consequences were in a wider system. Network calls, emails, and git commits are outside the sandbox and therefore outside the grade.
Embedded agent collapses status detail
The embedded agent records every failure as a single error status, losing the distinction between timeout, malformed response, and policy denial. The trajectory has the truth. The agent view does not.
Replay is single-run
Replay reproduces one recorded run. There is no batch replay and no side-by-side agent comparison yet. Both are planned, neither is here.
Fault-level statistics missing
The report shows pass rates by task, not by fault type. Adding fault-level statistics requires a small infrastructure change that has not been made yet.
The full list is in the repository README.
Takeaways
What this project changed in how I work.
A clean pass rate is a warning, not a result.
The first MVP passed fifty out of fifty. That was the signal to redesign, not to ship.
Evaluation frameworks need negative controls.
A grader that has never failed a known-bad agent has not been tested. The oracle matrix is not optional.
State-based grading sees what sequence-based grading cannot.
Two identical tool calls can mean success or duplication. Only the resulting state tells them apart.
End-to-end tests find what code review does not.
Three bugs had been sitting in the proxy for two phases. They only surfaced when a new fault mode exercised the path they lived on.
Deciding what not to build is part of the build.
Idempotency was deliberately delayed until the side-effect layer existed. The LLM judge was deferred to v0.3. The restraint is the reason the current version is coherent.