We study how AI systems improve after they reach production.
Research, field studies and practical methods for turning production failures into better agents.
The hard part is closing it.
Production → Traces → Diagnosis → Evaluation → Intervention
A deployed system produces evidence no pre-launch process can: traces, incidents, silent corrections, rejected outputs, and the workarounds users invent when the system is almost good enough.
The loop breaks at the conversion step. A failure gets noticed, a fix gets shipped, and a test gets written that encodes that one incident. The suite grows. Confidence grows. Neither tracks the system’s actual reliability.
Bring an agent that works 70–90% of the time and 100 traces.
Apply to the Eval Clinic