We study how AI systems improve after they reach production.

Research, field studies and practical methods for turning production failures into better agents.

Most teams have every input to this loop.

The hard part is closing it.

Production, traces, diagnosis, evaluation, intervention, and back to production. 01Production 02Traces 03Diagnosis 04Evaluation 05Intervention the loop closes here self-improving loop

Production → Traces → Diagnosis → Evaluation → Intervention

Launch is just a start. The most effective agents quickly improve in production.

A deployed system produces evidence no pre-launch process can: traces, incidents, silent corrections, rejected outputs, and the workarounds users invent when the system is almost good enough.

The loop breaks at the conversion step. A failure gets noticed, a fix gets shipped, and a test gets written that encodes that one incident. The suite grows. Confidence grows. Neither tracks the system’s actual reliability.

Fix your agent.

Bring an agent that works 70–90% of the time and 100 traces.

Apply to the Eval Clinic