Private suites · Public grades
What a coding‑agent subscription is actually worth
We run every coding agent on private task suites, on the plan you would actually buy. Two numbers per agent: how hard a problem it can take, and what a task costs in the plan's own quota.
Choosing a plan? How to choose · Two-minute quiz
Latest
All news- Platform UpdateNew suite: supervisor
- W37 grading runClaude Fable 5.1 beats out GPT-6 Astra - Week 37 agent results
- GPT-6 Astra releaseGPT-6 Astra: no better than Sol, worse than Fable
LeaderboardW37 2026
Last updated
| # | Agent | Overall | Plan usage | Approx. cost | Real-world issues | Spec planning | Vibe coding |
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1claude-code | C+ | ~12% of weekly limitestClaude Max 5x6% of weekly limitClaude Max 20xnot availableAnthropic does not offer Fable on Claude Pro - it ships only on the Max tiers, so there is no $20 plan to run this agent on and nothing to estimate.~12% of weekly limitestClaude Max 5x6% of weekly limitClaude Max 20x | ≈$0.145/taskest$0.145/tasknot availableAnthropic does not offer Fable on Claude Pro - it ships only on the Max tiers, so there is no $20 plan to run this agent on and nothing to estimate.≈$0.145/taskest$0.145/task | C+ | C+ | C+ |
| 2 | Claude Opus 5claude-code | C | ~4% of weekly limitestClaude Max 5x2% of weekly limitClaude Max 20x~40% of weekly limitestClaude Pro~4% of weekly limitestClaude Max 5x2% of weekly limitClaude Max 20x | ≈$0.048/taskest$0.048/task≈$0.097/taskest≈$0.048/taskest$0.048/task | D+ | C+ | C+ |
| 3 | GPT-6 Astracodex-cli | C | 5% of weekly limitChatGPT Pro 5x5% of weekly limitChatGPT Pro 5x~25% of weekly limitestChatGPT Plus5% of weekly limitChatGPT Pro 5x~1.3% of weekly limitestChatGPT Pro 20x | $0.061/task$0.061/task≈$0.061/taskest$0.061/task≈$0.030/taskest | D | B- | C+ |
| 4 | Grok 4.6grok-cli | C- | 8.8% of weekly limitSuperGrok8.8% of weekly limitSuperGrok8.8% of weekly limitSuperGrokno estimateno estimate | $0.032/task$0.032/task$0.032/task-- | C- | C | D+ |
| 5 | GPT-5.6 Solcodex-cli | C- | 3% of weekly limitChatGPT Pro 5x3% of weekly limitChatGPT Pro 5x~15% of weekly limitestChatGPT Plus3% of weekly limitChatGPT Pro 5x~0.8% of weekly limitestChatGPT Pro 20x | $0.036/task$0.036/task≈$0.036/taskest$0.036/task≈$0.018/taskest | C- | C | D+ |
| 6 | Gemini 3.8 Flashantigravity-cli | C- | 23.7% of weekly limitGoogle AI Pro23.7% of weekly limitGoogle AI Pro23.7% of weekly limitGoogle AI Pro~4.7% of weekly limitestGoogle AI Ultra 5x~1.2% of weekly limitestGoogle AI Ultra 20x | $0.057/task$0.057/task$0.057/task≈$0.046/taskest≈$0.026/taskest | C+ | D | C- |
| 7 | GPT-5.6 Terracodex-cli | D+ | 2% of weekly limitChatGPT Pro 5x2% of weekly limitChatGPT Pro 5x~10% of weekly limitestChatGPT Plus2% of weekly limitChatGPT Pro 5x~0.5% of weekly limitestChatGPT Pro 20x | $0.024/task$0.024/task≈$0.024/taskest$0.024/task≈$0.012/taskest | D+ | C | D |
| 8 | Kimi K3 256Kkimi-cli | D | 35% of weekly limitKimi Allegretto35% of weekly limitKimi Allegrettono estimateno estimateno estimate | $0.165/task$0.165/task--- | C- | D | D- |
| 9 | Composer 2.5cursor-cli | D | 2.1% of monthly limitCursor Pro2.1% of monthly limitCursor Pro2.1% of monthly limitCursor Pro~0.7% of monthly limitestCursor Pro+~0.1% of monthly limitestCursor Ultra | $0.022/task$0.022/task$0.022/task≈$0.022/taskest≈$0.011/taskest | C- | D | D- |
| 10 | GPT-5.6 Lunacodex-cli | D | <1% of weekly limitChatGPT Pro 5x<1% of weekly limitChatGPT Pro 5x~<5% of weekly limitestChatGPT Plus<1% of weekly limitChatGPT Pro 5x~<0.3% of weekly limitestChatGPT Pro 20x | <$0.012/task<$0.012/task≈<$0.012/taskest<$0.012/task≈<$0.006/taskest | D+ | D+ | F |
| 11 | Claude Sonnet 5claude-code | D | ~2% of weekly limitestClaude Max 5x1% of weekly limitClaude Max 20x~20% of weekly limitestClaude Pro~2% of weekly limitestClaude Max 5x1% of weekly limitClaude Max 20x | ≈$0.024/taskest$0.024/task≈$0.048/taskest≈$0.024/taskest$0.024/task | C- | D- | F |
| 12 | Claude Haiku 4.5claude-code | F | ~2% of weekly limitestClaude Max 5x1% of weekly limitClaude Max 20x~20% of weekly limitestClaude Pro~2% of weekly limitestClaude Max 5x1% of weekly limitClaude Max 20x | ≈$0.024/taskest$0.024/task≈$0.048/taskest≈$0.024/taskest$0.024/task | D+ | F | F |
| 13 | Gemini 3.1 Proantigravity-cli | F | 13.5% of weekly limitGoogle AI Pro13.5% of weekly limitGoogle AI Pro13.5% of weekly limitGoogle AI Pro~2.7% of weekly limitestGoogle AI Ultra 5x~0.7% of weekly limitestGoogle AI Ultra 20x | $0.033/task$0.033/task$0.033/task≈$0.026/taskest≈$0.015/taskest | D | F | F |
Default view: Claude rows are estimated at Claude Max 5x (~$100) from the Max 20x run they were measured on, marked est. The weekly limit grows ~2x from Max 5x to Max 20x, not the advertised session 4x. OpenAI rows are measured on ChatGPT Pro 5x. Every other row reads on the plan it ran on. Pick "Measured plan" for every row as measured, or a budget band to re-express all of them.
Budget view: rows whose measured plan sits in the picked band keep their measured figure; the rest are scaled from the measured plan by the provider's own advertised tier multiples (marked est) - approximate, never measured. Providers that publish no multiple get no estimate. Our paired Claude Max 5x/20x runs measured ~5.7x where 4x is advertised, so treat estimates as rough. A model the tier does not sell at all - Fable on Claude Pro - reads not available rather than an estimate.
Grades are within-suite-version comparable only. See methodology. Or compare two agents head-to-head. Overall is the capability mean across a subject's graded suite families.
How it works
Task suites stay private so they cannot leak into training data - what is public is the methodology: ladder shape, weights, scoring, and grade bands.
Each agent runs every suite headless through its own tools, with its harness version and plan pinned: capability, plan cost, task time, and reliability, once per task per weekly run.
Only the graded outcome is published. Third-party claims are labeled context, never blended into our grades.