Menu

Categories

Tags

Stretch your Claude Code tokens with Anthropic's 11 tips

August 15, 2026 | Source: t | Anthropic, Developer | 153 views 0 comments

Anthropic just published a guide to making Claude Code tokens stretch further, and its core lesson is a simple one: the longer your context, the more every turn costs. Chat history, file contents, and command output all keep eating tokens as the conversation goes on.

Here are the 11 most practical tips, straight from the source.

1. Switch tasks? /clear. When a task is done, clear the context. If you're still on the same task but the earlier conversation is no longer useful, use /compact to compress it. Long sessions get more expensive as they go, because every turn carries the full history.

2. Pick your model and effort from the start. Running /model or /effort mid-task invalidates the prompt cache, so the next turn has to reprocess the entire context. Same with Fast mode — turn it on from the beginning.

3. Run /compact before you step away. Cached prompts last about an hour for subscribers and default to roughly five minutes for API keys. If you return after the cache expires, the next turn usually re-reads the whole context. Compressing while the cache is still warm is cheaper.

4. Going off the rails? /rewind. /rewind only drops the last few turns, so the earlier cache stays valid. /compact rewrites the whole conversation, which costs more.

5. Attach files with @filename. Claude pulls the file in on the first turn, saving a Read or a search. A file usually only needs one @; calling it again can shove a duplicate copy into the context.

6. Be specific in your prompt. Instead of "tests are broken," tell Claude which test and which file. Otherwise it may grep, search, and read a pile of files — and all those results will sit in the context for every single turn after.

7. Keep test and log output lean. Build output, test output, and git log all stay in the context. Use quiet flags, print errors only, or trim with tail. Claude Code automatically writes command output over 30,000 characters to a file, leaving just the summary and path in the session.

8. Check startup overhead with /context. At the start of a session, /context shows how much CLAUDE.md, MCP servers, and tool definitions are taking up. Put only general rules in CLAUDE.md, move specialized workflows into on-demand Skills, and turn off unused MCPs with /mcp.

9. Use smaller models and lower effort for simple tasks. Bigger models cost more on input and output. Thinking tokens count as output tokens. For pure execution, you don't need the biggest model at maximum reasoning effort.

10. Leave big logs to subagents, but don't abuse them for small tasks. Subagents have their own context and only bring the final result back to the main session, making them great for reading large logs or searching docs. But they also re-read files and run their own multi-turn loops, so on tiny tasks they can end up costing more. For search and log reading, you can point them at Haiku or Sonnet.

11. Run /loop in its own session. Each loop is a full turn that carries the entire current context. If the interval also exceeds the cache lifetime, you get an extra cache miss on top. Anthropic recommends starting a clean terminal session just for loops.

The takeaway: stuff less junk into context, clear it when a task wraps, and avoid marathon sessions.

Read the full walkthrough from Anthropic: Anthropic

Tags: #Claude

Leave a Reply

Your email address will not be published. Required fields are marked *