Use fewer tokens in Claude Code
Make a Pro plan or a Team Standard seat last longer by seeing what uses your limit, changing a few habits, and trimming what Claude reads.
- Platform
- macOS on Apple silicon
- Time
- About 20 minutes
- Steps
- 14
On a Pro plan or a Team Standard seat, Claude Code draws from a usage allowance that resets every five hours and every week. Most of that allowance goes on Claude re-reading your conversation and your files, not on the answers it writes. So the biggest savings come from short conversations, a lighter model for everyday work, and less noise in what Claude reads. The first steps show where your usage goes, the middle steps change the habits and settings that matter most, and the last steps cover community tools and what to do when you run out.
Verified with Claude Code 2.1.283 · macOS 27.0
steps
Step 1: How your usage limit worksTwo windows, one allowance shared with Claude chat, and why a long conversation costs more with every message.
Your plan gives you an allowance with two limits: one that resets on a rolling five-hour window, and one that resets weekly. The same allowance covers Claude Code, Claude chat and Cowork, so a long chat in the browser leaves less for coding. On a Team plan, a Premium seat has about five times the usage of a Standard seat.
Claude Code does not send only your latest message. Every request carries the whole conversation so far, and Claude makes a new request each time it reads a file, runs a command or gets a tool result back. Prompt caching makes the repeated part cheaper, but it still counts, so a one-line question at the end of an all-day conversation costs as much as re-reading the whole day.
These are the things that make usage climb fastest:
What Why it costs Long conversations Every step re-sends everything before it. Long breaks On a subscription the cache lasts an hour after your last message. After a longer break, the next message processes the whole conversation again without the cache discount. Opus, high effort Opus costs more than Sonnet, and the thinking that higher effort adds is billed like the answer itself. Big tool output Test runs, logs and whole files stay in the conversation and are re-sent with every later step. Extra agents Subagents, agent teams and repeating /looptasks each send their own requests.The rest of this guide takes these on one by one.
Step 2: See what is using your limitThree built-in commands show how much you have used and what used it, and a status line keeps it in view.
Start by finding out where your usage goes, so you fix the right thing.
In Claude Code /usage/usageshows how much of your five-hour and weekly limits you have used. On a Pro or Team plan it also breaks down what used it: the share that went to skills, subagents, plugins and each MCP server, plus a warning for any habit, such as long conversations or cache misses, that caused 10% or more of your recent usage. Presswto see the last 7 days anddfor the last 24 hours. The breakdown only counts Claude Code on this computer, not Claude chat in the browser.In Claude Code /context/contextdraws a colored grid of what fills the current conversation: instructions, tools, memory files and messages. It also suggests what to trim.In Claude Code /insights/insightsreads your recent sessions and writes a report about how you work, where Claude misunderstood you, and what to try instead. The report itself uses tokens, so run it now and then, not every day.Keep an eye on it while you work. A status line is the bar under the prompt. This one shows the model and how full the conversation is, and, on Pro and Max plans, how much of the five-hour limit you have used. Claude Code does not send the limit to the status line on Team plans, so there it shows only the first two. It needs
jq:Terminal brew install jqThis writes the script, backing up any script with the same name first, and turns it on only if you do not already have a status line:
Terminal mkdir -p ~/.claude [ -f ~/.claude/statusline.sh ] && cp ~/.claude/statusline.sh ~/.claude/statusline.sh.bak-$(date +%Y%m%d-%H%M%S) cat > ~/.claude/statusline.sh <<'EOF' #!/bin/bash # Model, context use and, on Pro and Max, 5-hour limit use. jq -r ' "\(.model.display_name) · \(.context_window.used_percentage // 0 | floor)% context" + (if .rate_limits.five_hour.used_percentage then " · \(.rate_limits.five_hour.used_percentage | floor)% of 5-hour limit" else "" end) ' EOF chmod +x ~/.claude/statusline.sh f=~/.claude/settings.json [ -f "$f" ] || echo '{}' > "$f" if jq -e '.statusLine' "$f" > /dev/null; then echo "You already have a status line, so it was left alone." else cp "$f" "$f.bak-$(date +%Y%m%d-%H%M%S)" jq '.statusLine = {"type": "command", "command": "~/.claude/statusline.sh"}' "$f" > "$f.tmp" && mv "$f.tmp" "$f" fiThe status line appears after your next message. It reads something like
Sonnet 5 · 23% context · 41% of 5-hour limit.Step 3: Start fresh between tasksOne task per conversation, a handoff note when a task runs long, and a new conversation after a long break.
This is the habit that saves the most. When you finish one task and start something unrelated, clear the conversation so the old one stops riding along with every new message:
In Claude Code /clearThe old conversation is not lost.
/clearwith a name, such as/clear login-bug, labels it, and/resumebrings it back later.Prefer
/clearto/compact./compactsummarizes the conversation so you can keep going, but it reads the whole conversation to write that summary, so on a long one it is a large request in itself./clearcosts nothing. Use/compactonly when you are halfway through a task and need to keep its details, and tell it what to keep:In Claude Code /compact keep the failing test names and the files we changedCarry work over with a handoff note. When a task runs long, ask Claude to write down where things stand before you clear:
In Claude Code Write what we did, what is left, and the files involved to handoff.mdThen run
/clearand pick up from the note:In Claude Code Read handoff.md and continue with what is leftCome back to a fresh conversation after a long break. The cache lasts an hour after your last message. The first message after a longer break processes the whole conversation again without the cache discount, so for a big conversation, starting fresh from a handoff note is cheaper than carrying on. On Pro and Max, Claude Code offers to resume a large session from a summary for the same reason.
Ask side questions with
/btw. A quick question you do not need to keep, such as how a command works, can skip the conversation:In Claude Code /btw what does git rebase --onto do?Step 4: Pick the model and effort for the jobUse Sonnet for everyday coding, save Opus for hard problems, and choose both at the start of a conversation.
Since Claude Code 2.1.280, the default model on Pro and Team plans is Opus 5.5. Before that, Pro and Team Standard defaulted to Sonnet 5, so if your usage started running out sooner after an update, this is a likely reason. Sonnet handles most coding well and costs less, so make it your default and switch to Opus when a problem needs it, such as a tricky design decision or a bug you cannot pin down:
In Claude Code /model sonnet/modelsaves your choice for new sessions too. To use Opus for one session only, run/model, move to Opus, and presssinstead of Enter.Choose at the start, not halfway. Each model has its own cache, so switching models mid-conversation makes the next request read the whole conversation again without the cache discount. Pick the model when you start a task, or
/clearfirst.opusplanis a middle ground: Opus while you plan in plan mode, Sonnet while it carries out the plan. Every switch between planning and doing is a model switch, so it suits one plan followed by a long stretch of work, not constant back and forth.In Claude Code /model opusplanTurn effort down for routine work. Effort is how much Claude thinks before it answers, and that thinking counts like the answer does. The levels run
low,medium,high,xhighandmax. Opus 5.5 starts atmediumand the other models athigh. For routine edits, renames and small fixes,mediumorlowis usually enough:In Claude Code /effort mediumOn most models, changing effort mid-conversation also resets the cache, so set it when you start.
/effort autogoes back to the model's default.Put subagents on Sonnet too. Subagents use your session's model unless told otherwise, so an Opus session runs Opus subagents. These settings run every subagent on Sonnet, whatever the main conversation uses. The command backs up your settings file first and needs
jq, which the usage step installs:Terminal f=~/.claude/settings.json [ -f "$f" ] || echo '{}' > "$f" cp "$f" "$f.bak-$(date +%Y%m%d-%H%M%S)" jq '.env.CLAUDE_CODE_SUBAGENT_MODEL = "sonnet" | .env.CLAUDE_CODE_SUBAGENT_MODEL_FORCE = "1"' "$f" > "$f.tmp" && mv "$f.tmp" "$f"It takes effect in your next session. To undo it, delete those two lines from the
envblock in~/.claude/settings.json.Step 5: Ask precisely and stop earlyName the files, plan big changes first, and stop Claude the moment it heads the wrong way.
A vague request makes Claude search the whole project before it starts. A precise one lets it open the right file straight away.
Instead of Try Improve this codebase Add input validation to the login function in @src/auth.ts Fix the tests npm testfails in @src/cart.test.ts with "total is NaN". Find the cause and fix itWhy is this slow? The /orders page takes 4 seconds. Check the query in @src/orders/list.ts first Typing
@and a path attaches that file, so Claude does not have to go looking for it.Say it all in one message. Put related requests, the constraints and the expected result in one message instead of drip-feeding them, since every extra round re-sends the conversation. Say how Claude can check its work, such as a test to run or the output you expect, so it can catch its own mistakes instead of waiting for you to spot them.
Plan bigger changes first. Press
Shift+Tabuntil the status bar shows plan mode. Claude reads the code and proposes a plan without changing anything, so a wrong direction costs one plan instead of a pile of edits to undo.Stop early, and rewind instead of arguing. Press
Escas soon as Claude heads the wrong way. Then go back to before the mistake instead of correcting it in the conversation, since every correction stays in the conversation and is re-sent with every later step:In Claude Code /rewindPressing
Esctwice on an empty prompt opens the same menu. Pick the message to go back to, restore the code, the conversation or both, then send a clearer version of your request.Step 6: Keep CLAUDE.md shortEverything in CLAUDE.md is sent with every request, so keep it to what Claude cannot work out on its own.
CLAUDE.mdloads at the start of every session and rides along with every request after that. So do the files it imports with@path, such as@docs/style.md. Anthropic recommends keeping eachCLAUDE.mdunder 200 lines.Count the lines in your personal file and the current project's file:
Terminal wc -l ~/.claude/CLAUDE.md CLAUDE.md .claude/CLAUDE.md 2>/dev/nullKeep what Claude cannot work out by reading the code:
- The commands to build, test and run the project.
- Conventions that differ from the usual, such as "use pnpm, not npm".
- Traps, such as "the staging database is shared, never reset it".
Move everything else out:
- Step-by-step workflows, such as releasing or writing a migration, become skills. Only a skill's one-line description loads at the start, and the rest loads when the skill is used.
- Rules for one part of the code, such as the API folder, become rules with
paths, which load only when Claude opens a matching file. - Anything Claude can see in the code, such as the folder layout or the framework you use, can go.
/contextshows how much room your memory files take, so you can check the result.Step 7: Turn off tools you do not useEvery connected MCP server and installed skill adds to every session, so keep only the ones you use.
An MCP server connects Claude Code to another tool, such as GitHub, a database or a browser. Claude Code loads each server's tool names and instructions into every session, and a tool's full description once Claude first uses it. A handful of servers you never use still cost something in every conversation.
List your servers and switch off the ones you do not need:
In Claude Code /mcpOr turn one off by name, replacing
server-name:In Claude Code /mcp disable server-nameTo remove one for good, list them from the terminal and remove it by name:
Terminal claude mcp list claude mcp remove server-nameUse a command line tool when there is one. Tools such as
ghfor GitHub,aws,gcloudandsentry-cliadd nothing to your sessions until Claude runs them, because Claude already knows how to use a terminal. If a command line tool does the job, you can remove the MCP server that does the same thing.Prune skills and plugins too. Each skill's description sits in the conversation on every turn, whether Claude uses the skill or not, and plugins often bring several skills at once. This report shows what each skill costs, how often it runs, and which plugins you have not used lately:
In Claude Code /skill-doctorIt opens in the Stats tab of the
/pluginmanager, where you can turn off what you do not need.Step 8: Keep Claude out of big generated foldersDeny rules stop Claude from reading dependencies, build output and lock files, which are large and rarely help.
Folders like
node_modulesanddist, and files likepackage-lock.json, are huge and generated. When Claude opens one while searching, whatever it reads stays in the conversation and is re-sent with every later step. AReaddeny rule blocks them for Claude's file tools, for@mentions, and for shell commands that name the file, such ascatandhead. It is a guardrail against wasted reads, not a security boundary: a script that opens files itself can still read them.Run this in your project's root folder. It adds the rules to the project's
.claude/settings.json, backs the file up first, skips any rule that is already there, and keeps your existing rules in their order. Edit the list to match your stack before you run it:Terminal mkdir -p .claude f=.claude/settings.json [ -f "$f" ] || echo '{}' > "$f" cp "$f" "$f.bak-$(date +%Y%m%d-%H%M%S)" jq '(.permissions.deny // []) as $old | .permissions.deny = $old + ([ "Read(node_modules/**)", "Read(dist/**)", "Read(build/**)", "Read(coverage/**)", "Read(.next/**)", "Read(package-lock.json)", "Read(yarn.lock)", "Read(pnpm-lock.yaml)", "Read(*.min.js)" ] - $old)' "$f" > "$f.tmp" && mv "$f.tmp" "$f"A folder rule such as
Read(dist/**)blocks a folder with that name at any depth, and a file name such asRead(yarn.lock)matches that file anywhere in the project. Start a new session so the rules take effect. Commit.claude/settings.jsonso everyone on the project gets them, and leave the.bak-copy out of the commit.If Claude needs one of these files for a task, such as reading a library's source while debugging, delete that rule for the task and put it back afterwards.
Step 9: Show Claude only the failing testsA hook trims test runs down to the failures before Claude reads them, so a passing suite no longer fills the conversation.
A full test run can print thousands of lines, and all of it stays in the conversation. This hook, adapted from the Claude Code docs, runs before every shell command Claude runs. When the command starts with
npm test,pytestorgo test, it rewrites it to keep only the lines aroundFAIL,ERRORanderror:, at most 100 lines. When nothing matches, Claude sees a one-line note instead of an empty result.Write the script, backing up any script with the same name first:
Terminal mkdir -p ~/.claude/hooks [ -f ~/.claude/hooks/filter-test-output.sh ] && cp ~/.claude/hooks/filter-test-output.sh ~/.claude/hooks/filter-test-output.sh.bak-$(date +%Y%m%d-%H%M%S) cat > ~/.claude/hooks/filter-test-output.sh <<'EOF' #!/bin/bash # Keep only the failures from test runs before Claude reads them. input=$(cat) cmd=$(echo "$input" | jq -r '.tool_input.command') if [[ "$cmd" =~ ^(npm test|pytest|go test) ]]; then filtered_cmd="$cmd 2>&1 | { grep -A 5 -E '(FAIL|ERROR|error:)' || echo 'No FAIL or ERROR lines in the test output.'; } | head -100" echo "$input" | jq --arg filtered "$filtered_cmd" \ '{hookSpecificOutput: {hookEventName: "PreToolUse", permissionDecision: "allow", updatedInput: (.tool_input + {command: $filtered})}}' else echo "{}" fi EOF chmod +x ~/.claude/hooks/filter-test-output.shThen register it in
~/.claude/settings.json, backing the file up first and skipping it if it is already there:Terminal f=~/.claude/settings.json [ -f "$f" ] || echo '{}' > "$f" if ! grep -q 'filter-test-output.sh' "$f"; then cp "$f" "$f.bak-$(date +%Y%m%d-%H%M%S)" jq '.hooks.PreToolUse += [{"matcher": "Bash", "hooks": [{"type": "command", "command": "~/.claude/hooks/filter-test-output.sh"}]}]' "$f" > "$f.tmp" && mv "$f.tmp" "$f" fiRun
/hooksin Claude Code to check that it is listed underPreToolUse.A few things to know:
- The hook approves the test command it rewrites, so Claude runs it without asking you first.
- The exit code Claude sees is no longer the test runner's, so it judges the run by the lines it gets back.
- Change the pattern
^(npm test|pytest|go test)to match how your project runs its tests, such as^(pnpm test|cargo test). - rtk, in the community tools step, filters test runs and many other commands. Use this hook or rtk, not both.
Step 10: Add a code intelligence pluginA language server lets Claude jump straight to a definition instead of searching and opening several files.
Without help, Claude finds a function by searching for its name and opening the files that match. A code intelligence plugin connects Claude Code to your language's language server, the same engine your editor uses for "go to definition", so one lookup replaces that search. It also reports type errors after each edit, so Claude catches mistakes without running a full build.
The plugins come from Anthropic's official marketplace. This adds that marketplace if Claude Code has not added it yet, and does nothing if it is already there:
Terminal claude plugin marketplace add anthropics/claude-plugins-officialEach plugin needs its language server installed first, so run the pair of commands for your language.
TypeScript and JavaScript. The
@6matters: TypeScript 7 no longer includes thetsserverthat this language server runs, so it needs TypeScript 6 installed globally. That global copy also covers projects that use TypeScript 7 themselves.Terminal npm install -g typescript-language-server typescript@6 claude plugin install typescript-lsp@claude-plugins-officialPython:
Terminal npm install -g pyright claude plugin install pyright-lsp@claude-plugins-officialGo:
Terminal go install golang.org/x/tools/gopls@latest claude plugin install gopls-lsp@claude-plugins-officialRust:
Terminal rustup component add rust-analyzer claude plugin install rust-analyzer-lsp@claude-plugins-officialThe Claude Code docs list more languages, including Java, C and C++, C#, PHP, Ruby, Kotlin and Swift.
Start a new session to load the plugin. After Claude edits a file, the type errors in that file reach it with its next step. If Claude never reports type errors for that language, run
/pluginand open the Errors tab:Executable not found in $PATHmeans the shell you startedclaudefrom cannot find the language server.Step 12: Community tools: what they really saveoptionalrtk, headroom, ponytail and caveman each trim a different part of the bill, and their measured savings are smaller than the headlines.
These open-source tools each shrink a different part of what you pay for. Their headline numbers measure only that part, so the effect on your whole limit is smaller. The Claude Code guide has install steps for each one.
Tool What it trims Measured savings rtk Shell command output, such as git, tests and builds Up to 90% of that output. The project points out this is not 90% off your bill, because command output is only one part of what Claude reads. headroom Large tool results, such as logs and files 21% to 57% on the project's own benchmark scenarios, and very little on prose or output that is already compact. ponytail How much code Claude writes JetBrains ran 80 paired tasks: 10% lower cost, 11% faster, and no measurable change in quality. The project advertises 20% cheaper. caveman The length of Claude's replies JetBrains ran 86 paired coding tasks: 8.5% fewer output tokens, roughly 10% lower expected cost per task, and no detectable change in quality. The project advertises 65%, which holds for chat-style questions rather than coding sessions. ponytail and caveman only change what Claude writes, and in a coding session Claude reads far more than it writes. The example session in Anthropic's cost docs has 5.3 thousand output tokens next to 940 thousand tokens re-read from the cache. That is why the habits in the earlier steps, such as clearing between tasks and choosing Sonnet, usually save more than any of these tools.
If you add tools, start with one that trims what Claude reads, rtk or headroom, rather than both, because they overlap. Check the plugin and skill shares in
/usageafter a week, and remove what does not pay for itself.Step 13: When you hit the limitRead which limit you hit, let Claude Code wait and carry on, or pay for extra usage if your plan allows it.
The message tells you which limit you hit and when it resets.
- "You've hit your session limit" or "You've hit your weekly limit" covers every model, so switching models does not help.
- A limit that names one model, such as "You've hit your Opus limit", does not block other models, so
/model sonnetkeeps you working.
Let Claude Code wait and carry on. Claude Code can wait for the reset and then continue the task it was on. It often offers this on its own when you hit the limit, and you can open the options yourself:
In Claude Code /rate-limit-optionsPick Wait here, then continue automatically. Claude Code still asks for permissions as usual, so the task can pause on a question while you are away.
Pay for extra usage. Usage credits let you keep working past the limit at standard API prices:
In Claude Code /usage-creditsOn a Pro plan this opens your usage settings, where you can turn credits on and set a monthly spending cap. On a Team plan without billing access, it sends a request to your admins instead. While you run on credits, the cache lasts five minutes instead of an hour, so pauses cost more there than inside your plan.
If you hit it every week anyway, check
/usagefor the biggest share first. On a Team plan, a Premium seat gives about five times the usage, and admins can give one to just the people who need it.Step 14: Daily checklistThe whole guide on one screen, to keep next to your terminal.
Once:
- Make Sonnet your default with
/model sonnet. - Add the status line from the usage step.
- Trim
CLAUDE.mdto under 200 lines, moving workflows to skills and area rules to.claude/rules/. - Turn off MCP servers, skills and plugins you do not use, with
/mcpand/skill-doctor. - Block generated folders with
Readdeny rules in.claude/settings.json. - Install the code intelligence plugin for your language.
Every task:
- One task per conversation, with
/clearin between. - Name the files with
@pathand say how to check the result. - Plan big changes first:
Shift+Tabto plan mode. - Press
Escas soon as it goes wrong, then/rewind. - After an hour away, write a handoff note and
/clear. - Ask side questions with
/btw.
Every week:
- Run
/usageand presswto see what used your limit. - Remove the plugins, skills and MCP servers that take a big share without paying for themselves.
- Make Sonnet your default with
sources
- Claude Code: Manage costs effectively
- Claude Code: How Claude Code uses prompt caching
- Claude Code: Model configuration
- Claude Help Center: Usage limit best practices
- Claude Code: Cache lifetime
- Claude Help Center: What is the Team plan?
- Claude Code: Track your costs
- Claude Code: Customize your status line
- Claude Code: Manage context proactively
- Claude Code: Why usage climbs in a long session
- Claude Code: Switching models and the cache
- Claude Code: Choose a model for subagents
- Claude Code: Write specific prompts
- Claude Code: Work efficiently on complex tasks
- Claude Code: Checkpointing
- Claude Code: Move instructions from CLAUDE.md to skills
- Claude Code: How Claude remembers your project
- Claude Code: Reduce MCP server overhead
- Claude Code: Connect to tools with MCP
- Claude Code: Find unused skills
- Claude Code: Read and Edit permission rules
- Claude Code: Offload processing to hooks and skills
- Claude Code: Hooks reference
- Claude Code: Code intelligence plugins
- Claude Code: Agent team token costs
- Claude Code: Which TTL each request gets
- JetBrains: Ponytail skill for Claude Code, tested
- JetBrains: Speak to AI agents like cavemen to save tokens
- rtk: How savings work
- headroom on GitHub
- Claude Code: Wait for a usage limit to reset
- Claude Code: When a developer asks about a limit
- Claude Help Center: Extra usage for paid Claude plans