Skip to content
whoami

Use fewer tokens in Claude Code

Make a Pro plan or a Team Standard seat last longer by seeing what uses your limit, changing a few habits, and trimming what Claude reads.

Platform
macOS on Apple silicon
Time
About 20 minutes
Steps
14

On a Pro plan or a Team Standard seat, Claude Code draws from a usage allowance that resets every five hours and every week. Most of that allowance goes on Claude re-reading your conversation and your files, not on the answers it writes. So the biggest savings come from short conversations, a lighter model for everyday work, and less noise in what Claude reads. The first steps show where your usage goes, the middle steps change the habits and settings that matter most, and the last steps cover community tools and what to do when you run out.

Verified with Claude Code 2.1.283 · macOS 27.0

steps

  1. Step 1: How your usage limit worksTwo windows, one allowance shared with Claude chat, and why a long conversation costs more with every message.

    Your plan gives you an allowance with two limits: one that resets on a rolling five-hour window, and one that resets weekly. The same allowance covers Claude Code, Claude chat and Cowork, so a long chat in the browser leaves less for coding. On a Team plan, a Premium seat has about five times the usage of a Standard seat.

    Claude Code does not send only your latest message. Every request carries the whole conversation so far, and Claude makes a new request each time it reads a file, runs a command or gets a tool result back. Prompt caching makes the repeated part cheaper, but it still counts, so a one-line question at the end of an all-day conversation costs as much as re-reading the whole day.

    These are the things that make usage climb fastest:

    What Why it costs
    Long conversations Every step re-sends everything before it.
    Long breaks On a subscription the cache lasts an hour after your last message. After a longer break, the next message processes the whole conversation again without the cache discount.
    Opus, high effort Opus costs more than Sonnet, and the thinking that higher effort adds is billed like the answer itself.
    Big tool output Test runs, logs and whole files stay in the conversation and are re-sent with every later step.
    Extra agents Subagents, agent teams and repeating /loop tasks each send their own requests.

    The rest of this guide takes these on one by one.

  2. Step 2: See what is using your limitThree built-in commands show how much you have used and what used it, and a status line keeps it in view.

    Start by finding out where your usage goes, so you fix the right thing.

    In Claude Code
    /usage

    /usage shows how much of your five-hour and weekly limits you have used. On a Pro or Team plan it also breaks down what used it: the share that went to skills, subagents, plugins and each MCP server, plus a warning for any habit, such as long conversations or cache misses, that caused 10% or more of your recent usage. Press w to see the last 7 days and d for the last 24 hours. The breakdown only counts Claude Code on this computer, not Claude chat in the browser.

    In Claude Code
    /context

    /context draws a colored grid of what fills the current conversation: instructions, tools, memory files and messages. It also suggests what to trim.

    In Claude Code
    /insights

    /insights reads your recent sessions and writes a report about how you work, where Claude misunderstood you, and what to try instead. The report itself uses tokens, so run it now and then, not every day.

    Keep an eye on it while you work. A status line is the bar under the prompt. This one shows the model and how full the conversation is, and, on Pro and Max plans, how much of the five-hour limit you have used. Claude Code does not send the limit to the status line on Team plans, so there it shows only the first two. It needs jq:

    Terminal
    brew install jq

    This writes the script, backing up any script with the same name first, and turns it on only if you do not already have a status line:

    Terminal
    mkdir -p ~/.claude
    [ -f ~/.claude/statusline.sh ] && cp ~/.claude/statusline.sh ~/.claude/statusline.sh.bak-$(date +%Y%m%d-%H%M%S)
    cat > ~/.claude/statusline.sh <<'EOF'
    #!/bin/bash
    # Model, context use and, on Pro and Max, 5-hour limit use.
    jq -r '
      "\(.model.display_name) · \(.context_window.used_percentage // 0 | floor)% context"
      + (if .rate_limits.five_hour.used_percentage
         then " · \(.rate_limits.five_hour.used_percentage | floor)% of 5-hour limit"
         else "" end)
    '
    EOF
    chmod +x ~/.claude/statusline.sh
    f=~/.claude/settings.json
    [ -f "$f" ] || echo '{}' > "$f"
    if jq -e '.statusLine' "$f" > /dev/null; then
      echo "You already have a status line, so it was left alone."
    else
      cp "$f" "$f.bak-$(date +%Y%m%d-%H%M%S)"
      jq '.statusLine = {"type": "command", "command": "~/.claude/statusline.sh"}' "$f" > "$f.tmp" && mv "$f.tmp" "$f"
    fi

    The status line appears after your next message. It reads something like Sonnet 5 · 23% context · 41% of 5-hour limit.

  3. Step 3: Start fresh between tasksOne task per conversation, a handoff note when a task runs long, and a new conversation after a long break.

    This is the habit that saves the most. When you finish one task and start something unrelated, clear the conversation so the old one stops riding along with every new message:

    In Claude Code
    /clear

    The old conversation is not lost. /clear with a name, such as /clear login-bug, labels it, and /resume brings it back later.

    Prefer /clear to /compact. /compact summarizes the conversation so you can keep going, but it reads the whole conversation to write that summary, so on a long one it is a large request in itself. /clear costs nothing. Use /compact only when you are halfway through a task and need to keep its details, and tell it what to keep:

    In Claude Code
    /compact keep the failing test names and the files we changed

    Carry work over with a handoff note. When a task runs long, ask Claude to write down where things stand before you clear:

    In Claude Code
    Write what we did, what is left, and the files involved to handoff.md

    Then run /clear and pick up from the note:

    In Claude Code
    Read handoff.md and continue with what is left

    Come back to a fresh conversation after a long break. The cache lasts an hour after your last message. The first message after a longer break processes the whole conversation again without the cache discount, so for a big conversation, starting fresh from a handoff note is cheaper than carrying on. On Pro and Max, Claude Code offers to resume a large session from a summary for the same reason.

    Ask side questions with /btw. A quick question you do not need to keep, such as how a command works, can skip the conversation:

    In Claude Code
    /btw what does git rebase --onto do?
  4. Step 4: Pick the model and effort for the jobUse Sonnet for everyday coding, save Opus for hard problems, and choose both at the start of a conversation.

    Since Claude Code 2.1.280, the default model on Pro and Team plans is Opus 5.5. Before that, Pro and Team Standard defaulted to Sonnet 5, so if your usage started running out sooner after an update, this is a likely reason. Sonnet handles most coding well and costs less, so make it your default and switch to Opus when a problem needs it, such as a tricky design decision or a bug you cannot pin down:

    In Claude Code
    /model sonnet

    /model saves your choice for new sessions too. To use Opus for one session only, run /model, move to Opus, and press s instead of Enter.

    Choose at the start, not halfway. Each model has its own cache, so switching models mid-conversation makes the next request read the whole conversation again without the cache discount. Pick the model when you start a task, or /clear first.

    opusplan is a middle ground: Opus while you plan in plan mode, Sonnet while it carries out the plan. Every switch between planning and doing is a model switch, so it suits one plan followed by a long stretch of work, not constant back and forth.

    In Claude Code
    /model opusplan

    Turn effort down for routine work. Effort is how much Claude thinks before it answers, and that thinking counts like the answer does. The levels run low, medium, high, xhigh and max. Opus 5.5 starts at medium and the other models at high. For routine edits, renames and small fixes, medium or low is usually enough:

    In Claude Code
    /effort medium

    On most models, changing effort mid-conversation also resets the cache, so set it when you start. /effort auto goes back to the model's default.

    Put subagents on Sonnet too. Subagents use your session's model unless told otherwise, so an Opus session runs Opus subagents. These settings run every subagent on Sonnet, whatever the main conversation uses. The command backs up your settings file first and needs jq, which the usage step installs:

    Terminal
    f=~/.claude/settings.json
    [ -f "$f" ] || echo '{}' > "$f"
    cp "$f" "$f.bak-$(date +%Y%m%d-%H%M%S)"
    jq '.env.CLAUDE_CODE_SUBAGENT_MODEL = "sonnet" | .env.CLAUDE_CODE_SUBAGENT_MODEL_FORCE = "1"' "$f" > "$f.tmp" && mv "$f.tmp" "$f"

    It takes effect in your next session. To undo it, delete those two lines from the env block in ~/.claude/settings.json.

  5. Step 5: Ask precisely and stop earlyName the files, plan big changes first, and stop Claude the moment it heads the wrong way.

    A vague request makes Claude search the whole project before it starts. A precise one lets it open the right file straight away.

    Instead of Try
    Improve this codebase Add input validation to the login function in @src/auth.ts
    Fix the tests npm test fails in @src/cart.test.ts with "total is NaN". Find the cause and fix it
    Why is this slow? The /orders page takes 4 seconds. Check the query in @src/orders/list.ts first

    Typing @ and a path attaches that file, so Claude does not have to go looking for it.

    Say it all in one message. Put related requests, the constraints and the expected result in one message instead of drip-feeding them, since every extra round re-sends the conversation. Say how Claude can check its work, such as a test to run or the output you expect, so it can catch its own mistakes instead of waiting for you to spot them.

    Plan bigger changes first. Press Shift+Tab until the status bar shows plan mode. Claude reads the code and proposes a plan without changing anything, so a wrong direction costs one plan instead of a pile of edits to undo.

    Stop early, and rewind instead of arguing. Press Esc as soon as Claude heads the wrong way. Then go back to before the mistake instead of correcting it in the conversation, since every correction stays in the conversation and is re-sent with every later step:

    In Claude Code
    /rewind

    Pressing Esc twice on an empty prompt opens the same menu. Pick the message to go back to, restore the code, the conversation or both, then send a clearer version of your request.

  6. Step 6: Keep CLAUDE.md shortEverything in CLAUDE.md is sent with every request, so keep it to what Claude cannot work out on its own.

    CLAUDE.md loads at the start of every session and rides along with every request after that. So do the files it imports with @path, such as @docs/style.md. Anthropic recommends keeping each CLAUDE.md under 200 lines.

    Count the lines in your personal file and the current project's file:

    Terminal
    wc -l ~/.claude/CLAUDE.md CLAUDE.md .claude/CLAUDE.md 2>/dev/null

    Keep what Claude cannot work out by reading the code:

    • The commands to build, test and run the project.
    • Conventions that differ from the usual, such as "use pnpm, not npm".
    • Traps, such as "the staging database is shared, never reset it".

    Move everything else out:

    • Step-by-step workflows, such as releasing or writing a migration, become skills. Only a skill's one-line description loads at the start, and the rest loads when the skill is used.
    • Rules for one part of the code, such as the API folder, become rules with paths, which load only when Claude opens a matching file.
    • Anything Claude can see in the code, such as the folder layout or the framework you use, can go.

    /context shows how much room your memory files take, so you can check the result.

  7. Step 7: Turn off tools you do not useEvery connected MCP server and installed skill adds to every session, so keep only the ones you use.

    An MCP server connects Claude Code to another tool, such as GitHub, a database or a browser. Claude Code loads each server's tool names and instructions into every session, and a tool's full description once Claude first uses it. A handful of servers you never use still cost something in every conversation.

    List your servers and switch off the ones you do not need:

    In Claude Code
    /mcp

    Or turn one off by name, replacing server-name:

    In Claude Code
    /mcp disable server-name

    To remove one for good, list them from the terminal and remove it by name:

    Terminal
    claude mcp list
    claude mcp remove server-name

    Use a command line tool when there is one. Tools such as gh for GitHub, aws, gcloud and sentry-cli add nothing to your sessions until Claude runs them, because Claude already knows how to use a terminal. If a command line tool does the job, you can remove the MCP server that does the same thing.

    Prune skills and plugins too. Each skill's description sits in the conversation on every turn, whether Claude uses the skill or not, and plugins often bring several skills at once. This report shows what each skill costs, how often it runs, and which plugins you have not used lately:

    In Claude Code
    /skill-doctor

    It opens in the Stats tab of the /plugin manager, where you can turn off what you do not need.

  8. Step 8: Keep Claude out of big generated foldersDeny rules stop Claude from reading dependencies, build output and lock files, which are large and rarely help.

    Folders like node_modules and dist, and files like package-lock.json, are huge and generated. When Claude opens one while searching, whatever it reads stays in the conversation and is re-sent with every later step. A Read deny rule blocks them for Claude's file tools, for @ mentions, and for shell commands that name the file, such as cat and head. It is a guardrail against wasted reads, not a security boundary: a script that opens files itself can still read them.

    Run this in your project's root folder. It adds the rules to the project's .claude/settings.json, backs the file up first, skips any rule that is already there, and keeps your existing rules in their order. Edit the list to match your stack before you run it:

    Terminal
    mkdir -p .claude
    f=.claude/settings.json
    [ -f "$f" ] || echo '{}' > "$f"
    cp "$f" "$f.bak-$(date +%Y%m%d-%H%M%S)"
    jq '(.permissions.deny // []) as $old | .permissions.deny = $old + ([
      "Read(node_modules/**)",
      "Read(dist/**)",
      "Read(build/**)",
      "Read(coverage/**)",
      "Read(.next/**)",
      "Read(package-lock.json)",
      "Read(yarn.lock)",
      "Read(pnpm-lock.yaml)",
      "Read(*.min.js)"
    ] - $old)' "$f" > "$f.tmp" && mv "$f.tmp" "$f"

    A folder rule such as Read(dist/**) blocks a folder with that name at any depth, and a file name such as Read(yarn.lock) matches that file anywhere in the project. Start a new session so the rules take effect. Commit .claude/settings.json so everyone on the project gets them, and leave the .bak- copy out of the commit.

    If Claude needs one of these files for a task, such as reading a library's source while debugging, delete that rule for the task and put it back afterwards.

  9. Step 9: Show Claude only the failing testsA hook trims test runs down to the failures before Claude reads them, so a passing suite no longer fills the conversation.

    A full test run can print thousands of lines, and all of it stays in the conversation. This hook, adapted from the Claude Code docs, runs before every shell command Claude runs. When the command starts with npm test, pytest or go test, it rewrites it to keep only the lines around FAIL, ERROR and error:, at most 100 lines. When nothing matches, Claude sees a one-line note instead of an empty result.

    Write the script, backing up any script with the same name first:

    Terminal
    mkdir -p ~/.claude/hooks
    [ -f ~/.claude/hooks/filter-test-output.sh ] && cp ~/.claude/hooks/filter-test-output.sh ~/.claude/hooks/filter-test-output.sh.bak-$(date +%Y%m%d-%H%M%S)
    cat > ~/.claude/hooks/filter-test-output.sh <<'EOF'
    #!/bin/bash
    # Keep only the failures from test runs before Claude reads them.
    input=$(cat)
    cmd=$(echo "$input" | jq -r '.tool_input.command')
    
    if [[ "$cmd" =~ ^(npm test|pytest|go test) ]]; then
      filtered_cmd="$cmd 2>&1 | { grep -A 5 -E '(FAIL|ERROR|error:)' || echo 'No FAIL or ERROR lines in the test output.'; } | head -100"
      echo "$input" | jq --arg filtered "$filtered_cmd" \
        '{hookSpecificOutput: {hookEventName: "PreToolUse", permissionDecision: "allow", updatedInput: (.tool_input + {command: $filtered})}}'
    else
      echo "{}"
    fi
    EOF
    chmod +x ~/.claude/hooks/filter-test-output.sh

    Then register it in ~/.claude/settings.json, backing the file up first and skipping it if it is already there:

    Terminal
    f=~/.claude/settings.json
    [ -f "$f" ] || echo '{}' > "$f"
    if ! grep -q 'filter-test-output.sh' "$f"; then
      cp "$f" "$f.bak-$(date +%Y%m%d-%H%M%S)"
      jq '.hooks.PreToolUse += [{"matcher": "Bash", "hooks": [{"type": "command", "command": "~/.claude/hooks/filter-test-output.sh"}]}]' "$f" > "$f.tmp" && mv "$f.tmp" "$f"
    fi

    Run /hooks in Claude Code to check that it is listed under PreToolUse.

    A few things to know:

    • The hook approves the test command it rewrites, so Claude runs it without asking you first.
    • The exit code Claude sees is no longer the test runner's, so it judges the run by the lines it gets back.
    • Change the pattern ^(npm test|pytest|go test) to match how your project runs its tests, such as ^(pnpm test|cargo test).
    • rtk, in the community tools step, filters test runs and many other commands. Use this hook or rtk, not both.
  10. Step 10: Add a code intelligence pluginA language server lets Claude jump straight to a definition instead of searching and opening several files.

    Without help, Claude finds a function by searching for its name and opening the files that match. A code intelligence plugin connects Claude Code to your language's language server, the same engine your editor uses for "go to definition", so one lookup replaces that search. It also reports type errors after each edit, so Claude catches mistakes without running a full build.

    The plugins come from Anthropic's official marketplace. This adds that marketplace if Claude Code has not added it yet, and does nothing if it is already there:

    Terminal
    claude plugin marketplace add anthropics/claude-plugins-official

    Each plugin needs its language server installed first, so run the pair of commands for your language.

    TypeScript and JavaScript. The @6 matters: TypeScript 7 no longer includes the tsserver that this language server runs, so it needs TypeScript 6 installed globally. That global copy also covers projects that use TypeScript 7 themselves.

    Terminal
    npm install -g typescript-language-server typescript@6
    claude plugin install typescript-lsp@claude-plugins-official

    Python:

    Terminal
    npm install -g pyright
    claude plugin install pyright-lsp@claude-plugins-official

    Go:

    Terminal
    go install golang.org/x/tools/gopls@latest
    claude plugin install gopls-lsp@claude-plugins-official

    Rust:

    Terminal
    rustup component add rust-analyzer
    claude plugin install rust-analyzer-lsp@claude-plugins-official

    The Claude Code docs list more languages, including Java, C and C++, C#, PHP, Ruby, Kotlin and Swift.

    Start a new session to load the plugin. After Claude edits a file, the type errors in that file reach it with its next step. If Claude never reports type errors for that language, run /plugin and open the Errors tab: Executable not found in $PATH means the shell you started claude from cannot find the language server.

  11. Step 11: Watch the hidden spendersSubagents, agent teams, repeating tasks and big pastes all use your limit in the background.

    Some features do useful work out of sight, and each one draws on the same limit as your conversation.

    What What to know
    Subagents Each one works in its own conversation and sends its own requests. Their cache lasts five minutes, not an hour. They are worth it for noisy jobs, such as digging through logs, because only a summary comes back, but asking for five in parallel for a small job costs five times over.
    Agent teams Several full Claude Code sessions working together. Anthropic measured about 7 times the tokens of a normal session when teammates run in plan mode. They are off unless you turn them on, and on a basic plan it is best to leave them off.
    Repeating tasks /loop and scheduled tasks send your whole conversation on every run, even while you are away. Stop them when you are done, and check the Loops rows in /usage.
    ultracode effort Plans a multi-agent workflow for each substantial task, which multiplies usage. Keep it for rare, large jobs.
    Big pastes A whole log or file pasted into the prompt stays in the conversation. Paste the error and a few lines around it, or give the file path.
    Fast mode /fast does not use your plan's limit at all: on subscription plans it bills usage credits, which cost money.

    If /usage shows subagents taking a large share, give the subagents you write a smaller model in their settings, such as model: haiku for simple lookups, or put them all on Sonnet as in the model step.

  12. Step 12: Community tools: what they really saveoptionalrtk, headroom, ponytail and caveman each trim a different part of the bill, and their measured savings are smaller than the headlines.

    These open-source tools each shrink a different part of what you pay for. Their headline numbers measure only that part, so the effect on your whole limit is smaller. The Claude Code guide has install steps for each one.

    Tool What it trims Measured savings
    rtk Shell command output, such as git, tests and builds Up to 90% of that output. The project points out this is not 90% off your bill, because command output is only one part of what Claude reads.
    headroom Large tool results, such as logs and files 21% to 57% on the project's own benchmark scenarios, and very little on prose or output that is already compact.
    ponytail How much code Claude writes JetBrains ran 80 paired tasks: 10% lower cost, 11% faster, and no measurable change in quality. The project advertises 20% cheaper.
    caveman The length of Claude's replies JetBrains ran 86 paired coding tasks: 8.5% fewer output tokens, roughly 10% lower expected cost per task, and no detectable change in quality. The project advertises 65%, which holds for chat-style questions rather than coding sessions.

    ponytail and caveman only change what Claude writes, and in a coding session Claude reads far more than it writes. The example session in Anthropic's cost docs has 5.3 thousand output tokens next to 940 thousand tokens re-read from the cache. That is why the habits in the earlier steps, such as clearing between tasks and choosing Sonnet, usually save more than any of these tools.

    If you add tools, start with one that trims what Claude reads, rtk or headroom, rather than both, because they overlap. Check the plugin and skill shares in /usage after a week, and remove what does not pay for itself.

  13. Step 13: When you hit the limitRead which limit you hit, let Claude Code wait and carry on, or pay for extra usage if your plan allows it.

    The message tells you which limit you hit and when it resets.

    • "You've hit your session limit" or "You've hit your weekly limit" covers every model, so switching models does not help.
    • A limit that names one model, such as "You've hit your Opus limit", does not block other models, so /model sonnet keeps you working.

    Let Claude Code wait and carry on. Claude Code can wait for the reset and then continue the task it was on. It often offers this on its own when you hit the limit, and you can open the options yourself:

    In Claude Code
    /rate-limit-options

    Pick Wait here, then continue automatically. Claude Code still asks for permissions as usual, so the task can pause on a question while you are away.

    Pay for extra usage. Usage credits let you keep working past the limit at standard API prices:

    In Claude Code
    /usage-credits

    On a Pro plan this opens your usage settings, where you can turn credits on and set a monthly spending cap. On a Team plan without billing access, it sends a request to your admins instead. While you run on credits, the cache lasts five minutes instead of an hour, so pauses cost more there than inside your plan.

    If you hit it every week anyway, check /usage for the biggest share first. On a Team plan, a Premium seat gives about five times the usage, and admins can give one to just the people who need it.

  14. Step 14: Daily checklistThe whole guide on one screen, to keep next to your terminal.

    Once:

    • Make Sonnet your default with /model sonnet.
    • Add the status line from the usage step.
    • Trim CLAUDE.md to under 200 lines, moving workflows to skills and area rules to .claude/rules/.
    • Turn off MCP servers, skills and plugins you do not use, with /mcp and /skill-doctor.
    • Block generated folders with Read deny rules in .claude/settings.json.
    • Install the code intelligence plugin for your language.

    Every task:

    • One task per conversation, with /clear in between.
    • Name the files with @path and say how to check the result.
    • Plan big changes first: Shift+Tab to plan mode.
    • Press Esc as soon as it goes wrong, then /rewind.
    • After an hour away, write a handoff note and /clear.
    • Ask side questions with /btw.

    Every week:

    • Run /usage and press w to see what used your limit.
    • Remove the plugins, skills and MCP servers that take a big share without paying for themselves.

sources