findings

What our
experiments show.

Agents that start with API context are less likely to make the same mistakes again.

Two experiments got Sannr to where it is today. For each one: what we expected, what we ran, what we found, and the benchmark we hold the product to.

accumulation

Notes
carried forward.

The benchmark

1.06incidents a session, cold (no context)
0.06incidents a session, warm (with context)

The number every new build has to match: one trap hit in 16 warm sessions.

Hypothesis
If each agent inherits the notes the agents before it wrote about an API, later agents will stop repeating their mistakes.
What we ran
40 sessions in four independent lanes, five generations each, with two arms side by side: cold (no context), starting from nothing every generation, and warm (with context), inheriting the notes its own lane had written. Generation 1 was set aside as calibration by a rule fixed in advance, leaving 16 comparable pairs. One small service, one planted, undocumented trap. We counted incidents: hits on the trap, plus calls to routes the contract never declared.
What we found
Cold agents averaged 1.06 incidents a session. Warm agents averaged 0.06: one trap hit in 16 sessions. Both arms finished every task, 100 percent.
Conclusion
An agent that inherits what the agents before it learned about an API rarely repeats their mistakes. Finishing the task was never the difference. How often they hit the same trap on the way there was.

changing API

Context through
an API change.

The benchmark

81persisted incidents, cold (no context)
6persisted incidents, warm (with context)

Across 18 sessions per arm after the change. The bar for keeping context current through a change.

Hypothesis
When an API changes, agents that carry their context through the change will repeat fewer mistakes than agents that start over.
What we ran
Two arms through an API change: cold (no context) and warm (with context), its context kept up to date through the change. 18 sessions in each arm after the change. We counted persisted incidents.
What we found
Cold agents ended with 81 persisted incidents. Warm agents ended with 6.
Conclusion
Agents that start with API context are less likely to make the same mistakes again, even after the API changes.