One Hung API Call Used to Kill My 1,000-Run Benchmark. Here's the Fix.
The experiment was fine. The runner was the bug. A single hung API call discarded hours of completed work, and I kept blaming the data. If your long-running eval keeps dying at 80%, this is the five-p






