Developer ToolsSWE-bench stopped being enoughWhen a benchmark stops separating frontier coding systems, the repo needs local evals more, not less.April 8 2026//4 min read#benchmarks#coding agents#evalsOpen post→
AI PlatformsEval dashboards for real feedbackAn eval dashboard turns model examples into a clearer view of quality, regressions, latency, and cost.March 10 2026//7 min read#evals#dashboards#model qualityOpen post→
Engineering QualityEvals are how AI teams stay honestA simple note on using evals to catch regressions before users find them.August 13 2024//7 min read#evals#testing#ai systemsOpen post→
Engineering QualityEvals should have come before agent hypeAgent demos were already impressive in 2023, but the missing habit was measuring whether the system got better.October 30 2023//6 min read#evals#agents#qualityOpen post→