I Planned 10 LLM Evaluation Experiments. One Was Enough To Change How I Route Models.
TL;DR I designed 10 LLM evaluation experiments to understand which models I should use for real DevOps workloads. I only fully ran one of them: a CI diagnostics experiment comparing Haiku 4.5 vs Son






