Eval infra note: all 98 clinic-ops tasks pass a no-op/random/hardcode/gold/idempotent QC battery. The models are the secondary story.
Reddit r/MLOps1w4 min read
disclosure: Co-creator, and company i work at sells RL environments. This post is about the grader though, since that's usually the part that quietly lies to you. The result I actually trust. v1 saturated: frontier models cleared 100 tasks with mean rewards from 0.76 to 1.00. All 98 v2 tasks now pass a battery where no-op fails, random actions fail, a hardcoded guess fails, gold passes, gold_alt passes, and the verifier is idempotent. Report is in QC_REPORT.md. What that looks like in practice: - Only the final Postgres state gets graded. The transcript isn't an input. - The verifier runs on t
