← posts / ai integrations

Effort levels in Claude Code: when max effort pays off and when it just burns tokens

Anthropic's effort deep dive (Terminal-Bench 3.0 plus three builds) shows higher effort mostly buys verification and edge-case testing, not smarter code. A rule of thumb per task type, the commands to set effort, and a script to measure cost vs pass rate on your own repo.

if.codesOct 1, 2026 · 9 min read#claude-code#ai-coding#llm#developer-toolsAI-assisted

Claude Code's effort setting changes how much compute the model puts into a task. Anthropic just published a deep dive on it, "Using Claude Code: Spending your effort" by Thariq Shihipar. It covers Terminal-Bench 3.0 runs plus three hand-built apps at different levels. The short version: extra effort mostly buys verification and edge-case testing, not smarter code. Below is a rule of thumb for each task type, the commands to switch levels, and a script to measure the trade-off on your own repo.

What "effort" actually is

The post calls effort an approximation of how much compute you want Claude to spend on a task. Its analogy is a deadline: you'd work differently with 12 hours than with 1. In practice, higher effort means Claude takes more independent action for judgement and verification.

Effort levels act like a deadline, changing how the model allocates its compute time.
Effort levels act like a deadline, changing how the model allocates its compute time.

The model configuration docs list five levels: low, medium, high, xhigh and max. Older models such as Opus 4.6 and Sonnet 4.6 have no xhigh. If you pick a level the model doesn't support, Claude Code falls back to the highest supported level below it. Defaults vary by model:

  • Opus 5.5 and Sonnet 5.5: medium
  • Opus 4.7: xhigh
  • Every other model that supports effort: high

The docs also say the scale is calibrated per model, so high on one model isn't the same amount of work as high on another.

What the data says

Scores and tokens both go up

On Terminal-Bench 3.0 (70 tasks, with the 4 GPU tasks excluded), every step up in effort raised both the benchmark score and the tokens used, for Opus 5.5 and Fable 5.1 alike. The cost side is large. The post reports Fable 5.1 using a median of 73k tokens per attempt at low and 222k at max, about 3x.

Effort fixes missed edge cases, not wrong approaches

Anthropic went through 370 Fable 5.1 attempts:

Higher effort levels often focus on the intricate 'gears' of edge-case verification.
Higher effort levels often focus on the intricate 'gears' of edge-case verification.
  • Low effort: 140 passed, 59 failed on missed edge cases
  • Max effort: 214 passed, 24 failed on missed edge cases

The post is clear about the limit. More effort cuts failures from missed edge cases, but it does not fix runs where the model took the wrong approach. If Claude misunderstands the problem, max effort just does the wrong thing more thoroughly.

Gains depend a lot on the domain

Pass rates from low effort to top effort, by category:

  • Security: 64% → 87%
  • Hardware: 34% → 75%
  • ML: 54% → 73%
  • Science: 41% → 61%
  • Software: 43% → 56%
  • Media: 18% → 30%
  • Operations: 12% → 22%

Operations moved the least. The chart labels this "rulebook-style work stays low": when a task depends on knowing the right procedure, more thinking doesn't help much.

What the extra tokens were spent on

The case studies show where the extra effort went:

  • html-js-filter (security): Fable 5.1 went from 1/5 at low to 5/5 at xhigh. The high-effort run reviewed its own draft adversarially, read the parser's source, wrote XSS test suites and built a random-document fuzzer.
  • mvcc-lsm-compaction (storage bug): at low effort Claude edited the code without building or testing it. At xhigh it reproduced the crash, wrote randomized tests and verified the fix.
  • cli-2ph-simplex (LP solver): low effort was a single-pass implementation with little testing. High effort checked results against a separate solver, found performance problems and reworked the algorithm.
  • gsea-proteomics: low effort used one data-prep method. High effort tried two and checked that the results changed the way they should.

In all four, the extra effort went into testing and checking the work, not into writing more clever code.

The three builds: wall-clock time

  • Underspecified fitness app: 1.5 min (low), 4 min (medium), 11 min (high), 67 min (max). Low produced a basic log with a graph. Max produced a much bigger app with heat charts and more features.
  • Config-menu redesign: 1 min (low) vs 28 min (max). Low produced an interactive sketch. Max produced a polished mockup with flow walkthroughs.
  • Highly specified implementation: 16 / 22 / 33 / 79 min (low / medium / high / max). With a detailed spec, the post says the levels "behaved much more similarly."

The takeaway: a vague prompt at max effort gets you Claude's own interpretation of the task, built out in full. A clear spec narrows the gap between levels, so paying for max buys less.

The gap between effort levels narrows when a task has a highly detailed specification.
The gap between effort levels narrows when a task has a highly detailed specification.

Put the three builds next to the benchmark and the pattern is the same. Past the level where the task is actually solved, extra minutes go into scope and verification. On a vague prompt that means features you didn't ask for; on a clear spec, more checking of work that was mostly right already. So the question per task is not "how hard is this?" but "how much of the outcome depends on edge cases that only testing finds?" That is what the rules below sort by.

Rule of thumb by task type

This follows the blog's recommendations and the "Choose an effort level" table in the docs:

  • Brainstorming, first sketches, renames, small edits → low. You review every result anyway, and a starting point that arrives sooner is worth more than polish.
  • Day-to-day feature work with a clear scope → medium. This is the default on Opus 5.5. The docs say Anthropic's testing found Opus 5.5 at medium matches or beats Opus 5 at high on coding evals, so don't copy your old Opus 5 setting over.
  • Bug fixes in existing code, anything with hidden edge cases → high. This is where the reproduce-test-verify behaviour pays for itself.
  • Security review, parsers, sanitizers, numerical or hardware-adjacent code → xhigh or max. These categories gained the most in the benchmark.
  • Long autonomous runs on hard problems you won't supervise → max. The docs warn that max "may show diminishing returns and is prone to overthinking," so test it before using it everywhere.
  • Procedure-heavy ops work → don't expect much from higher effort. Better context (runbooks, CLAUDE.md) is likely a better investment. That's our reading of the ops numbers, not a claim the post makes.

The post also suggests a workflow for features that splits effort across phases:

  1. Ask Claude to interview you about missing details.
  2. Implement at low effort.
  3. Review and iterate at low effort.
  4. Verify and test at high effort.
Want AI wired into the systems you already run?I build LLM integrations with costs and quality you can see. The estimate is free.

How to set effort

All of these come from the model config docs and CLI reference:

# Inside a session: open the slider, or set a level directly
/effort
/effort high
/effort auto        # clear the saved level for the active model

# One session only, at launch (does not persist)
claude --effort xhigh

# Environment variable (also the only way to make max stick)
export CLAUDE_CODE_EFFORT_LEVEL=high

Some details worth knowing:

  • Typing a level after /effort, or pressing Enter in the slider, saves it as your default for that model. Pressing s in the slider applies it to the current session only.
  • max is session-only unless you set it with CLAUDE_CODE_EFFORT_LEVEL. The effortLevel setting accepts low through xhigh, not max.
  • You can set effort in skill or subagent frontmatter, so a security-review subagent can run at xhigh while your main session stays at medium.
  • Putting ultrathink in a prompt asks for deeper reasoning on that turn only, without changing your session level.
  • Resolution order: explicit choice (env var, --effort, /effort), then your settings, then the model default.

Measure it on your own repo

Benchmarks average over tasks that aren't yours, so the useful number is cost against pass rate on your own code. The script below runs the same prompt once per level, each in a throwaway git worktree. It records cost, duration, turns and output tokens from the --output-format json result, then runs your test command to see whether the change actually works.

Warning: this script runs Claude unattended with --permission-mode acceptEdits and the Bash tool allowed, so it can edit files and run any shell command in the worktree without asking. Run it only on a repo you trust, ideally in a container or VM, and narrow --allowedTools to your test command (for example "Bash(npm test *)") before pointing it at anything important. --max-budget-usd caps the spend per run.

#!/usr/bin/env bash
# effort-bench.sh - run one task at each effort level and compare.
# Usage: ./effort-bench.sh "Fix the flaky date parsing in src/dates.ts" "npm ci && npm test"
# Requires: git, jq, claude. Run from the repo root with a clean HEAD.
set -uo pipefail

PROMPT="${1:?usage: $0 \"task prompt\" \"test command\"}"
TEST_CMD="${2:?usage: $0 \"task prompt\" \"test command\"}"
LEVELS="${LEVELS:-low medium high xhigh max}"
BUDGET="${BUDGET:-10}"          # USD cap per run
OUT="$PWD/effort-results.jsonl"
LOGDIR="$PWD/effort-logs"       # stderr of each run, for debugging
mkdir -p "$LOGDIR"
: > "$OUT"

for level in $LEVELS; do
  wt="$(mktemp -d)/effort-$level"
  git worktree add --detach "$wt" HEAD >/dev/null 2>&1

  # With --output-format json, stdout carries only the final result object
  # (failures inside the run included); warnings go to stderr, kept in a log.
  json=$(cd "$wt" && claude -p "$PROMPT" \
    --effort "$level" \
    --output-format json \
    --permission-mode acceptEdits \
    --allowedTools "Read,Edit,Write,Bash" \
    --max-budget-usd "$BUDGET" \
    --no-session-persistence 2>"$LOGDIR/$level.log")
  [ -n "$json" ] || json='{"subtype":"no_result"}'   # bad flag or crash: see the log

  if (cd "$wt" && bash -c "$TEST_CMD" >/dev/null 2>&1); then pass=true; else pass=false; fi

  echo "$json" | jq -c --arg level "$level" --argjson pass "$pass" '{
    level: $level,
    tests_pass: $pass,
    cost_usd: .total_cost_usd,
    minutes: (((.duration_ms // 0) / 60000) * 10 | round / 10),
    turns: .num_turns,
    output_tokens: .usage.output_tokens,
    subtype: .subtype
  }' | tee -a "$OUT"

  git worktree remove --force "$wt"
done

echo
jq -s -r '.[] | "\(.level)\tpass=\(.tests_pass)\t$\(.cost_usd)\t\(.minutes)m\tturns=\(.turns)"' "$OUT"

Notes before you run it:

  • --allowedTools "Bash" lets Claude run any shell command in the worktree without asking. Only use it on a repo you trust, or narrow it with rule syntax such as "Bash(npm test *)".
  • total_cost_usd is a client-side estimate, not your bill. The docs say to use the Console or the Usage and Cost API for real billing numbers.
  • usage leaves out subagent tokens. For whole-tree accounting, read modelUsage.
  • Each worktree is a fresh checkout without node_modules, virtualenvs or build output, so make the test command install what it needs (as in "npm ci && npm test"), or every level will fail for the same boring reason.
  • If a row shows subtype other than success (for example error_max_budget_usd or no_result), check effort-logs/<level>.log before you compare numbers.
  • One run per level is noisy. Run each level three times and pick tasks you already know the answer to: a bug you fixed last month, or a feature with an existing test suite.

What to look for: the lowest level where tests pass consistently. If high passes and max only adds minutes and dollars, make high your default for that kind of task.

Sources

if.codesI build RAG, AI integrations and agent pipelines on Go and Python backends — and write about it here.
// keep reading

More posts on AI and backends.

// free quote

Read something you need? I’ll quote it for free.

RAG, AI integrations, agents or the backend underneath — tell me what you have and what should change. I read every request myself.

Free · no commitment

Tell me what you have. I’ll tell you what it takes.

1Describe the projectA few sentences is enough — about two minutes.
2I review itI read it myself and may ask a follow-up question.
3You get a free quoteScope, approach and estimate — yours to keep, no strings.
What kind of project is it?
Free and without obligation. Your details are used only to reply — see the privacy policy.