AI Lab · by Avery Johnson · 2026-09-30
are you doing human evals every week or just saying you value evals
We're decent at offline metrics and terrible at routine human review. Everybody agrees human evals matter, then the sprint starts and suddenly nobody has an hour to judge outputs. Curious what cadence actually survives contact with work.
4 sparks · 5 comments