Claude Code's effort setting changes how much compute the model puts into a task. Anthropic just published a deep dive on it, "Using Claude Code: Spending your effort" by Thariq Shihipar. It covers Terminal-Bench 3.0 runs plus three hand-built apps at different levels. The short version: extra effort mostly buys verification and edge-case testing, not smarter code. Below is a rule of thumb for each task type, the commands to switch levels, and a script to measure the trade-off on your own repo.
What "effort" actually is
The post calls effort an approximation of how much compute you want Claude to spend on a task. Its analogy is a deadline: you'd work differently with 12 hours than with 1. In practice, higher effort means Claude takes more independent action for judgement and verification.

The model configuration docs list five levels: low, medium, high, xhigh and max. Older models such as Opus 4.6 and Sonnet 4.6 have no xhigh. If you pick a level the model doesn't support, Claude Code falls back to the highest supported level below it. Defaults vary by model:
- Opus 5.5 and Sonnet 5.5:
medium - Opus 4.7:
xhigh - Every other model that supports effort:
high
The docs also say the scale is calibrated per model, so high on one model isn't the same amount of work as high on another.
What the data says
Scores and tokens both go up
On Terminal-Bench 3.0 (70 tasks, with the 4 GPU tasks excluded), every step up in effort raised both the benchmark score and the tokens used, for Opus 5.5 and Fable 5.1 alike. The cost side is large. The post reports Fable 5.1 using a median of 73k tokens per attempt at low and 222k at max, about 3x.
Effort fixes missed edge cases, not wrong approaches
Anthropic went through 370 Fable 5.1 attempts:

- Low effort: 140 passed, 59 failed on missed edge cases
- Max effort: 214 passed, 24 failed on missed edge cases
The post is clear about the limit. More effort cuts failures from missed edge cases, but it does not fix runs where the model took the wrong approach. If Claude misunderstands the problem, max effort just does the wrong thing more thoroughly.
Gains depend a lot on the domain
Pass rates from low effort to top effort, by category:
- Security: 64% → 87%
- Hardware: 34% → 75%
- ML: 54% → 73%
- Science: 41% → 61%
- Software: 43% → 56%
- Media: 18% → 30%
- Operations: 12% → 22%
Operations moved the least. The chart labels this "rulebook-style work stays low": when a task depends on knowing the right procedure, more thinking doesn't help much.
What the extra tokens were spent on
The case studies show where the extra effort went:
html-js-filter(security): Fable 5.1 went from 1/5 at low to 5/5 at xhigh. The high-effort run reviewed its own draft adversarially, read the parser's source, wrote XSS test suites and built a random-document fuzzer.mvcc-lsm-compaction(storage bug): at low effort Claude edited the code without building or testing it. At xhigh it reproduced the crash, wrote randomized tests and verified the fix.cli-2ph-simplex(LP solver): low effort was a single-pass implementation with little testing. High effort checked results against a separate solver, found performance problems and reworked the algorithm.gsea-proteomics: low effort used one data-prep method. High effort tried two and checked that the results changed the way they should.
In all four, the extra effort went into testing and checking the work, not into writing more clever code.
The three builds: wall-clock time
- Underspecified fitness app: 1.5 min (low), 4 min (medium), 11 min (high), 67 min (max). Low produced a basic log with a graph. Max produced a much bigger app with heat charts and more features.
- Config-menu redesign: 1 min (low) vs 28 min (max). Low produced an interactive sketch. Max produced a polished mockup with flow walkthroughs.
- Highly specified implementation: 16 / 22 / 33 / 79 min (low / medium / high / max). With a detailed spec, the post says the levels "behaved much more similarly."
The takeaway: a vague prompt at max effort gets you Claude's own interpretation of the task, built out in full. A clear spec narrows the gap between levels, so paying for max buys less.

Put the three builds next to the benchmark and the pattern is the same. Past the level where the task is actually solved, extra minutes go into scope and verification. On a vague prompt that means features you didn't ask for; on a clear spec, more checking of work that was mostly right already. So the question per task is not "how hard is this?" but "how much of the outcome depends on edge cases that only testing finds?" That is what the rules below sort by.
Rule of thumb by task type
This follows the blog's recommendations and the "Choose an effort level" table in the docs:
- Brainstorming, first sketches, renames, small edits →
low. You review every result anyway, and a starting point that arrives sooner is worth more than polish. - Day-to-day feature work with a clear scope →
medium. This is the default on Opus 5.5. The docs say Anthropic's testing found Opus 5.5 at medium matches or beats Opus 5 at high on coding evals, so don't copy your old Opus 5 setting over. - Bug fixes in existing code, anything with hidden edge cases →
high. This is where the reproduce-test-verify behaviour pays for itself. - Security review, parsers, sanitizers, numerical or hardware-adjacent code →
xhighormax. These categories gained the most in the benchmark. - Long autonomous runs on hard problems you won't supervise →
max. The docs warn that max "may show diminishing returns and is prone to overthinking," so test it before using it everywhere. - Procedure-heavy ops work → don't expect much from higher effort. Better context (runbooks, CLAUDE.md) is likely a better investment. That's our reading of the ops numbers, not a claim the post makes.
The post also suggests a workflow for features that splits effort across phases:
- Ask Claude to interview you about missing details.
- Implement at low effort.
- Review and iterate at low effort.
- Verify and test at high effort.


