Electricity Bench

Private suites · Public grades

What a coding‑agent subscription is actually worth

We run every coding agent on private task suites, on the plan you would actually buy. Two numbers per agent: how hard a problem it can take, and what a task costs in the plan's own quota.

Latest

All news

LeaderboardW37 2026

Last updated

Scroll horizontally to view all leaderboard columns.
#AgentOverallPlan usageApprox. costReal-world issuesSpec planningVibe coding
1Claude Fable 5.1claude-codeC+~12% of weekly limitestClaude Max 5x6% of weekly limitClaude Max 20xnot availableAnthropic does not offer Fable on Claude Pro - it ships only on the Max tiers, so there is no $20 plan to run this agent on and nothing to estimate.~12% of weekly limitestClaude Max 5x6% of weekly limitClaude Max 20x≈$0.145/taskest$0.145/tasknot availableAnthropic does not offer Fable on Claude Pro - it ships only on the Max tiers, so there is no $20 plan to run this agent on and nothing to estimate.≈$0.145/taskest$0.145/taskC+C+C+
2Claude Opus 5claude-codeC~4% of weekly limitestClaude Max 5x2% of weekly limitClaude Max 20x~40% of weekly limitestClaude Pro~4% of weekly limitestClaude Max 5x2% of weekly limitClaude Max 20x≈$0.048/taskest$0.048/task≈$0.097/taskest≈$0.048/taskest$0.048/taskD+C+C+
3GPT-6 Astracodex-cliC5% of weekly limitChatGPT Pro 5x5% of weekly limitChatGPT Pro 5x~25% of weekly limitestChatGPT Plus5% of weekly limitChatGPT Pro 5x~1.3% of weekly limitestChatGPT Pro 20x$0.061/task$0.061/task≈$0.061/taskest$0.061/task≈$0.030/taskestDB-C+
4Grok 4.6grok-cliC-8.8% of weekly limitSuperGrok8.8% of weekly limitSuperGrok8.8% of weekly limitSuperGrokno estimateno estimate$0.032/task$0.032/task$0.032/task--C-CD+
5GPT-5.6 Solcodex-cliC-3% of weekly limitChatGPT Pro 5x3% of weekly limitChatGPT Pro 5x~15% of weekly limitestChatGPT Plus3% of weekly limitChatGPT Pro 5x~0.8% of weekly limitestChatGPT Pro 20x$0.036/task$0.036/task≈$0.036/taskest$0.036/task≈$0.018/taskestC-CD+
6Gemini 3.8 Flashantigravity-cliC-23.7% of weekly limitGoogle AI Pro23.7% of weekly limitGoogle AI Pro23.7% of weekly limitGoogle AI Pro~4.7% of weekly limitestGoogle AI Ultra 5x~1.2% of weekly limitestGoogle AI Ultra 20x$0.057/task$0.057/task$0.057/task≈$0.046/taskest≈$0.026/taskestC+DC-
7GPT-5.6 Terracodex-cliD+2% of weekly limitChatGPT Pro 5x2% of weekly limitChatGPT Pro 5x~10% of weekly limitestChatGPT Plus2% of weekly limitChatGPT Pro 5x~0.5% of weekly limitestChatGPT Pro 20x$0.024/task$0.024/task≈$0.024/taskest$0.024/task≈$0.012/taskestD+CD
8Kimi K3 256Kkimi-cliD35% of weekly limitKimi Allegretto35% of weekly limitKimi Allegrettono estimateno estimateno estimate$0.165/task$0.165/task---C-DD-
9Composer 2.5cursor-cliD2.1% of monthly limitCursor Pro2.1% of monthly limitCursor Pro2.1% of monthly limitCursor Pro~0.7% of monthly limitestCursor Pro+~0.1% of monthly limitestCursor Ultra$0.022/task$0.022/task$0.022/task≈$0.022/taskest≈$0.011/taskestC-DD-
10GPT-5.6 Lunacodex-cliD<1% of weekly limitChatGPT Pro 5x<1% of weekly limitChatGPT Pro 5x~<5% of weekly limitestChatGPT Plus<1% of weekly limitChatGPT Pro 5x~<0.3% of weekly limitestChatGPT Pro 20x<$0.012/task<$0.012/task≈<$0.012/taskest<$0.012/task≈<$0.006/taskestD+D+F
11Claude Sonnet 5claude-codeD~2% of weekly limitestClaude Max 5x1% of weekly limitClaude Max 20x~20% of weekly limitestClaude Pro~2% of weekly limitestClaude Max 5x1% of weekly limitClaude Max 20x≈$0.024/taskest$0.024/task≈$0.048/taskest≈$0.024/taskest$0.024/taskC-D-F
12Claude Haiku 4.5claude-codeF~2% of weekly limitestClaude Max 5x1% of weekly limitClaude Max 20x~20% of weekly limitestClaude Pro~2% of weekly limitestClaude Max 5x1% of weekly limitClaude Max 20x≈$0.024/taskest$0.024/task≈$0.048/taskest≈$0.024/taskest$0.024/taskD+FF
13Gemini 3.1 Proantigravity-cliF13.5% of weekly limitGoogle AI Pro13.5% of weekly limitGoogle AI Pro13.5% of weekly limitGoogle AI Pro~2.7% of weekly limitestGoogle AI Ultra 5x~0.7% of weekly limitestGoogle AI Ultra 20x$0.033/task$0.033/task$0.033/task≈$0.026/taskest≈$0.015/taskestDFF

Default view: Claude rows are estimated at Claude Max 5x (~$100) from the Max 20x run they were measured on, marked est. The weekly limit grows ~2x from Max 5x to Max 20x, not the advertised session 4x. OpenAI rows are measured on ChatGPT Pro 5x. Every other row reads on the plan it ran on. Pick "Measured plan" for every row as measured, or a budget band to re-express all of them.

Grades are within-suite-version comparable only. See methodology. Or compare two agents head-to-head. Overall is the capability mean across a subject's graded suite families.

How it works

1
Private suites

Task suites stay private so they cannot leak into training data - what is public is the methodology: ladder shape, weights, scoring, and grade bands.

2
Graded runs

Each agent runs every suite headless through its own tools, with its harness version and plan pinned: capability, plan cost, task time, and reliability, once per task per weekly run.

3
Public scorecards

Only the graded outcome is published. Third-party claims are labeled context, never blended into our grades.