- Hypothesis
- If each agent inherits the notes the agents before it wrote about an API, later agents will stop repeating their mistakes.
- What we ran
- 40 sessions in four independent lanes, five generations each, with two arms side by side: cold (no context), starting from nothing every generation, and warm (with context), inheriting the notes its own lane had written. Generation 1 was set aside as calibration by a rule fixed in advance, leaving 16 comparable pairs. One small service, one planted, undocumented trap. We counted incidents: hits on the trap, plus calls to routes the contract never declared.
- What we found
- Cold agents averaged 1.06 incidents a session. Warm agents averaged 0.06: one trap hit in 16 sessions. Both arms finished every task, 100 percent.
- Conclusion
- An agent that inherits what the agents before it learned about an API rarely repeats their mistakes. Finishing the task was never the difference. How often they hit the same trap on the way there was.