Back to Portfolio
MVP — v0.2Open SourceMIT License

MCP-Faultline

Break your MCP agent. Then prove it recovered.

An external evaluator for MCP agents. Faultline injects controlled failures into tool calls, records the trajectory, and grades whether the agent recovered. With evidence, not just pass/fail.

18
Fault Types
5
Grading Layers
21/21
Oracle Validation
CI
Green on Linux

The Problem

Most agent benchmarks measure capability. Almost none measure resilience.

Two agents can both pass a task and behave completely differently when a tool times out. One retries once and recovers. Another retries fifty times and loops. A third declares success while the workspace is empty. Standard benchmarks would score all three identically.

The gap is not in the agent. It is in the evaluation. A pass/fail checker sees two successful tool calls. It does not see that the second call wrote a duplicate file because the first one already succeeded before its response timed out.

The red flag that started this project

My first MVP ran 50 executions and passed all 50. A clean 100% pass rate. That was not a signal of quality. It was a signal that the framework had never proven it could catch a failure.

The Pivot

Not more features. A different question.

I shared the MVP with five AI systems for review. All five converged on the same criticism, phrased different ways. The sharpest one came from ChatGPT:

Do not prove your agent is good. Prove Faultline can prove when your agent is bad.

That reframed the whole project. The goal was no longer to add features. It was to make the evaluator itself falsifiable. Every subsequent decision followed from that.

The architecture changed from an embedded harness, where Faultline ran the agent and therefore knew too much, to an external stdio proxy. The agent now runs outside, on its own. Faultline sits in the middle and watches traffic. A real LLM-driven agent, Cline, was integrated as the first external test subject.

The Solution

External proxy. Five-layer grading. State over sequence.

The agent is a black box. Faultline provides tools, injects faults at the protocol boundary, records every call, and grades the result.

      Task (YAML)
          |
          v
   +--------------+
   |   Sandbox    |   isolated temp folder
   +--------------+
          |
          v
   +--------------+       +---------------+
   |    Agent     | <---> |  Fault Layer  |   inject / forward
   +--------------+       +---------------+
          |                       |
          |                       v
          |              +------------------+
          |              |  Real MCP Server |
          |              +------------------+
          v
   +--------------+
   |  Trajectory  |   every call, every result
   +--------------+
          |
          v
   +--------------+
   |   Grader     |   5 layers
   +--------------+
          |
          v
      Report (JSON / Markdown)

Five-layer grading

Layer 1
Outcome

Did the final workspace match the task?

Layer 2
Safety

Did the agent stay within allowed tools and constraints?

Layer 3
Recovery

What strategy did the agent use, and did it work?

Layer 4
Integrity

What side effects actually happened?

Layer 5 — informational
Efficiency

How many calls, retries, duplicates? Metadata, not pass/fail.

Detail

18 fault types. Six categories. One registry.

The early version had four faults. Timeout, error, malformed, partial. That is a demo, not infrastructure. The current version has eighteen, grouped by the layer of the protocol they attack.

CategoryFaults
Transporttimeout, latency, disconnect, rate_limit
Protocolmalformed_response, invalid_schema, missing_field
Dataempty, partial_result, stale, contradictory, wrong_value
Authorizationauth_expired, forbidden
Temporaldelayed, out_of_order, replayed
Systemicburst_failure, server_outage, cascading_failure

The registry pattern exists because of one practical problem. Four faults spread across three switch statements is manageable. Eighteen faults across the same three would be fifty-four places to update, and any forgotten one would crash at runtime. The registry is one file. Adding a fault means adding one entry.

Idempotency violations get their own mode: mode: after. The tool runs first, then the response is replaced. A write succeeds on disk, the agent sees a timeout, retries, and writes twice. A sequence-based grader would see two successful calls. A state-based grader sees two side effects where there should be one.

A failure the grader actually caught

$ npm run test:idempotency-smoke
...
✅ integrity assertion fails
   File written 2 times (expected exactly 1): output.txt
✅ file was actually written (side-effect happened)

Total: 5/5 passed

This is the case that the sequence-based approach structurally cannot see. The agent did nothing wrong from its own view. It called the tool, got a timeout, and retried. The problem is that the first call already succeeded before the response was replaced.

Finding this required fixing three bugs that had been dormant since the proxy phase. A counter that never advanced, a call number that stuck at one, and an integrity check that trusted the status field instead of the executed flag. End-to-end testing caught all three. Code review did not.

Validation

21/21. Because the grader has to be able to fail.

A framework that catches nothing has proven nothing. To validate the grader itself, I built five control agents: one good, four broken in different ways. The matrix is 21 entries. Good agents must pass. Bad agents must fail, in the specific way that matches their defect.

Good Agent

Retries once, recovers, produces correct state.

Bad Retry Agent

Retries fifty times. Graded as looped.

Infinite Loop Agent

Calls the same tool forever. Graded as looped.

Unsafe Agent

Falls back to a forbidden tool. Graded as unsafe-recovery.

Hallucinating Agent

Claims success while the workspace is empty. Graded as false-recovery.

Total: 21/21 passed
✅ ALL CONTROLS PASSED — grader can distinguish good from bad agents.

Method

Five reviewers, before writing a line of production code.

Before the pivot, I sent the MVP and its planning documents to five different AI systems. Each one reviewed independently. The goal was not validation. It was to find what I had missed. All five converged on the same structural criticism, which is what made it credible.

The habit stuck. For each major phase since, I have run the same loop: write the plan, share it with several models, synthesize what is useful, discard what is not, and document the decision. The framework in this repository is the result of that process, not a single draft.

You are testing a puppet whose strings you hold.

One of the five, on the early embedded-harness design.

Limitations

Four limitations I chose not to hide.

Intent, not consequence

The current mode grades what the agent attempted. Not what the consequences were in a wider system. Network calls, emails, and git commits are outside the sandbox and therefore outside the grade.

Embedded agent collapses status detail

The embedded agent records every failure as a single error status, losing the distinction between timeout, malformed response, and policy denial. The trajectory has the truth. The agent view does not.

Replay is single-run

Replay reproduces one recorded run. There is no batch replay and no side-by-side agent comparison yet. Both are planned, neither is here.

Fault-level statistics missing

The report shows pass rates by task, not by fault type. Adding fault-level statistics requires a small infrastructure change that has not been made yet.

The full list is in the repository README.

Takeaways

What this project changed in how I work.

1

A clean pass rate is a warning, not a result.

The first MVP passed fifty out of fifty. That was the signal to redesign, not to ship.

2

Evaluation frameworks need negative controls.

A grader that has never failed a known-bad agent has not been tested. The oracle matrix is not optional.

3

State-based grading sees what sequence-based grading cannot.

Two identical tool calls can mean success or duplication. Only the resulting state tells them apart.

4

End-to-end tests find what code review does not.

Three bugs had been sitting in the proxy for two phases. They only surfaced when a new fault mode exercised the path they lived on.

5

Deciding what not to build is part of the build.

Idempotency was deliberately delayed until the side-effect layer existed. The LLM judge was deferred to v0.3. The restraint is the reason the current version is coherent.

Read the code. Run the tests. Break it.

The repository is open. If you find a way to make the evaluator miss a failure, that is the most useful feedback the project can get.