Articles & Snippets
Cursor Agent Models: Pricing, Scores, and How to Choose
Posted by negraru on Tue, 22 Sep 2026
A practical guide to model cost, capability, and when to pick Composer, Grok, Opus, Fable, or budget options.
Why this article?
Cursor bills agent and chat usage in tokens, not flat “requests.” Models sit in two usage pools that reset monthly. Capability varies widely, and the model picker’s default (often Composer Fast or Grok) is not always the best fit for small daily tasks or for hard multi-file agent runs.
This article summarizes public data from Cursor Models & Pricing and CursorBench 4.0, plus composite scores on a 0–100 scale so you can compare options at a glance. Scores are interpretive, not an official Cursor rating.
How Cursor pricing works?
Cursor Models pool
Grok 4.7, Grok 4.6, Grok 4.5, and Composer 2.5. Pro, Pro Plus, and Ultra include more included usage in this pool than for third-party models. No Cursor Token surcharge on Teams for these first-party models.
Other Models pool
Anthropic, OpenAI, Google, Meta, and others at list API rates. On Teams and Enterprise, third-party usage also incurs a $0.25 per million tokens Cursor Token Rate (individual Pro+ plans use the published list prices against your included allowance).
All prices below are USD per million tokens unless noted. Fast tiers (Grok, Composer) charge higher per-token rates for lower latency; they do not change underlying model intelligence. Effort levels (Grok, Opus, Fable) use the same list price but consume more tokens on harder settings.
The 0–100 score (methodology)
100 is a theoretical ceiling—“best realistic coding agent on the hardest public evals today”—not a model you can select. No shipping model sits at 100; even CursorBench 4.0 leaders fail a large share of tasks.
Scores are anchored to CursorBench 4.0: ambiguous, multi-file tasks from real Cursor sessions. Reference points:
- Claude Opus 5.5 Max — 57.8% on CB 4.0 → about 93 on this scale
- Composer 2.5 — 27.7% on CB 4.0 → about 84 (but lowest cost per task on that benchmark)
Other benchmarks (SWE-Bench Multilingual, Terminal-Bench 2.0 vs 4.0) measure different things; Composer can look stronger on older SWE-style numbers and weaker on CB 4.0’s long-horizon tasks. Treat the score as agent coding under stress, not “CSS tweak quality.”
Head-to-head: models you asked about
CB 4.0 = CursorBench 4.0 pass rate. ~$/task = average cost per task on CB 4.0 (Cursor-reported). Fast Grok rows use the same CB score; cost scales ~2× for the same tokens.
| Model | Pool | Score | CB 4.0 | ~$/task | Input | Output |
|---|---|---|---|---|---|---|
| Claude Opus 5.5 Medium | Other | 91 | 52.5% | $2.91 | $4 | $20 |
| Claude Fable 5.1 High | Other | 90 | 49.2% | $9.08 | $10 | $50 |
| Claude Opus 5 High | Other | 89 | 44.7% | $9.00 | $5 | $25 |
| Grok 4.7 High | Cursor | 88 | 43.9% | $4.69 | $2 | $6 |
| Grok 4.7 High Fast | Cursor | 88 | — | ~2× | $4 | $12 |
| Grok 4.6 High | Cursor | 88 | 40.4% | $5.20 | $2 | $6 |
| Grok 4.6 High Fast | Cursor | 88 | — | ~2× | $4 | $12 |
| Gemini 3.8 Flash High | Other | 88 | 39.6% | $4.70 | $0.75 | $3.50 |
| Muse Spark 1.3 High | Other | 86 | 33.4% | $1.66 | $1.25 | $4.25 |
| GPT-5.6 Sol Medium | Other | 85 | 31.1% | $1.77 | $4 | $20 |
| Composer 2.5 (standard) | Cursor | 84 | 27.7% | $0.68 | $0.50 | $2.50 |
Takeaways: Opus 5.5 Medium leads on score-per-dollar for hard agent tasks. Composer 2.5 is the default “daily driver” for cost. Grok 4.7 beats 4.6 slightly at the same price; Fast Grok is for latency, not IQ. Fable 5.1 High scores well but is expensive per task.
Cheapest models (token list price)
On paid Cursor plans, the lowest per-million-token third-party models include GPT-5.6 Luna and GPT-5.4 Nano (about $0.20 input / ~$1.20–1.25 output). In the Cursor Models pool, Composer 2.5 standard (not Fast) is the economical agent pick ($0.50 / $2.50). Hobby is the only free tier—limited Auto usage, no Cloud Agents.
Local open models vs Composer (composite scores)
These are not in Cursor’s picker by default; scores compare overall coding assistant capability, not Cursor integration.
| Model | Score (0–100) | Notes |
|---|---|---|
| Composer 2.5 | 84 | Native Cursor agent; best value in Cursor Models pool |
| Qwen3-Coder-30B-A3B | 73 | Strong open MoE coder; SWE scores depend heavily on agent harness |
| Qwen2.5-Coder-14B | 58 | Good small local model; not comparable to frontier agents |
Who is “100”? What about Claude Mythos?
Claude Mythos 5.1 is the same underlying model as Claude Fable 5.1 with different safeguards and invite-only access (cyber / life sciences programs)—not a higher tier for everyday coding.
The closest public peak for hard agent coding in Cursor evals is Claude Opus 5.5 at maximum effort (~57.8% CursorBench 4.0, roughly 93/100 on this scale). Fable 5.1 Max sits around 90–91; Grok 4.7 Extra High around 88–89.
Recommendations
Pro+ and small front-end work
- Default: Composer 2.5 standard (avoid Composer Fast unless you need speed).
- Stuck on a nasty bug or large refactor: Grok 4.7 High or Claude Opus 5.5 Medium.
- Must maximize pass rate: Opus 5.5 at higher effort or Fable 5.1 Max—watch cost.
- Preserve Other Models budget: Do routine work on Composer; use Luna only for high-volume cheap loops if enabled in settings.
Avoid surprises
- Pin a model instead of Auto if you want predictable spend (Auto routes to capable, pricier models on hard prompts).
- Check Dashboard → Spending for pool usage.
- Composer 2.5 (Fast) and Grok Fast cost multiples of standard tiers.
Sources & caveats
- Cursor — Models & Pricing
- Cursor — CursorBench 4.0
- Cursor — Composer 2.5 benchmarks
- Anthropic — Claude Opus 5.5
- Anthropic — Fable & Mythos 5.1
Benchmarks differ by version, effort, and agent harness. Small score gaps may not be statistically meaningful. Prices and model lists change; verify on Cursor before budgeting.
Cursor Agent Models Pricing Scores