Articles & Snippets

Cursor Agent Models: Pricing, Scores, and How to Choose

A practical guide to model cost, capability, and when to pick Composer, Grok, Opus, Fable, or budget options.

Why this article?

Cursor bills agent and chat usage in tokens, not flat “requests.” Models sit in two usage pools that reset monthly. Capability varies widely, and the model picker’s default (often Composer Fast or Grok) is not always the best fit for small daily tasks or for hard multi-file agent runs.

This article summarizes public data from Cursor Models & Pricing and CursorBench 4.0, plus composite scores on a 0–100 scale so you can compare options at a glance. Scores are interpretive, not an official Cursor rating.

How Cursor pricing works?

Cursor Models pool

Grok 4.7, Grok 4.6, Grok 4.5, and Composer 2.5. Pro, Pro Plus, and Ultra include more included usage in this pool than for third-party models. No Cursor Token surcharge on Teams for these first-party models.

Other Models pool

Anthropic, OpenAI, Google, Meta, and others at list API rates. On Teams and Enterprise, third-party usage also incurs a $0.25 per million tokens Cursor Token Rate (individual Pro+ plans use the published list prices against your included allowance).

All prices below are USD per million tokens unless noted. Fast tiers (Grok, Composer) charge higher per-token rates for lower latency; they do not change underlying model intelligence. Effort levels (Grok, Opus, Fable) use the same list price but consume more tokens on harder settings.

The 0–100 score (methodology)

100 is a theoretical ceiling—“best realistic coding agent on the hardest public evals today”—not a model you can select. No shipping model sits at 100; even CursorBench 4.0 leaders fail a large share of tasks.

Scores are anchored to CursorBench 4.0: ambiguous, multi-file tasks from real Cursor sessions. Reference points:

  • Claude Opus 5.5 Max — 57.8% on CB 4.0 → about 93 on this scale
  • Composer 2.5 — 27.7% on CB 4.0 → about 84 (but lowest cost per task on that benchmark)

Other benchmarks (SWE-Bench Multilingual, Terminal-Bench 2.0 vs 4.0) measure different things; Composer can look stronger on older SWE-style numbers and weaker on CB 4.0’s long-horizon tasks. Treat the score as agent coding under stress, not “CSS tweak quality.”

Head-to-head: models you asked about

CB 4.0 = CursorBench 4.0 pass rate. ~$/task = average cost per task on CB 4.0 (Cursor-reported). Fast Grok rows use the same CB score; cost scales ~2× for the same tokens.

Model Pool Score CB 4.0 ~$/task Input Output
Claude Opus 5.5 Medium Other 91 52.5% $2.91 $4 $20
Claude Fable 5.1 High Other 90 49.2% $9.08 $10 $50
Claude Opus 5 High Other 89 44.7% $9.00 $5 $25
Grok 4.7 High Cursor 88 43.9% $4.69 $2 $6
Grok 4.7 High Fast Cursor 88 ~2× $4 $12
Grok 4.6 High Cursor 88 40.4% $5.20 $2 $6
Grok 4.6 High Fast Cursor 88 ~2× $4 $12
Gemini 3.8 Flash High Other 88 39.6% $4.70 $0.75 $3.50
Muse Spark 1.3 High Other 86 33.4% $1.66 $1.25 $4.25
GPT-5.6 Sol Medium Other 85 31.1% $1.77 $4 $20
Composer 2.5 (standard) Cursor 84 27.7% $0.68 $0.50 $2.50

Takeaways: Opus 5.5 Medium leads on score-per-dollar for hard agent tasks. Composer 2.5 is the default “daily driver” for cost. Grok 4.7 beats 4.6 slightly at the same price; Fast Grok is for latency, not IQ. Fable 5.1 High scores well but is expensive per task.

Cheapest models (token list price)

On paid Cursor plans, the lowest per-million-token third-party models include GPT-5.6 Luna and GPT-5.4 Nano (about $0.20 input / ~$1.20–1.25 output). In the Cursor Models pool, Composer 2.5 standard (not Fast) is the economical agent pick ($0.50 / $2.50). Hobby is the only free tier—limited Auto usage, no Cloud Agents.

Local open models vs Composer (composite scores)

These are not in Cursor’s picker by default; scores compare overall coding assistant capability, not Cursor integration.

Model Score (0–100) Notes
Composer 2.5 84 Native Cursor agent; best value in Cursor Models pool
Qwen3-Coder-30B-A3B 73 Strong open MoE coder; SWE scores depend heavily on agent harness
Qwen2.5-Coder-14B 58 Good small local model; not comparable to frontier agents

Who is “100”? What about Claude Mythos?

Claude Mythos 5.1 is the same underlying model as Claude Fable 5.1 with different safeguards and invite-only access (cyber / life sciences programs)—not a higher tier for everyday coding.

The closest public peak for hard agent coding in Cursor evals is Claude Opus 5.5 at maximum effort (~57.8% CursorBench 4.0, roughly 93/100 on this scale). Fable 5.1 Max sits around 90–91; Grok 4.7 Extra High around 88–89.

Recommendations

Pro+ and small front-end work

  • Default: Composer 2.5 standard (avoid Composer Fast unless you need speed).
  • Stuck on a nasty bug or large refactor: Grok 4.7 High or Claude Opus 5.5 Medium.
  • Must maximize pass rate: Opus 5.5 at higher effort or Fable 5.1 Max—watch cost.
  • Preserve Other Models budget: Do routine work on Composer; use Luna only for high-volume cheap loops if enabled in settings.

Avoid surprises

  • Pin a model instead of Auto if you want predictable spend (Auto routes to capable, pricier models on hard prompts).
  • Check Dashboard → Spending for pool usage.
  • Composer 2.5 (Fast) and Grok Fast cost multiples of standard tiers.

Sources & caveats

Benchmarks differ by version, effort, and agent harness. Small score gaps may not be statistically meaningful. Prices and model lists change; verify on Cursor before budgeting.

Compiled from public documentation and leaderboard data, September 2026. Not affiliated with Cursor or model providers.


Cursor Agent  Models  Pricing  Scores