Hitting session and usage limits in Claude Code is rarely caused by the length of the prompts you write—actual user input accounts for as little as 0.01% of total token spend. Because large language models lack persistent memory, every subsequent prompt resends the entire accumulated conversation history. Understanding the mechanics of context accumulation, prompt caching, and background tooling is essential to cutting operational costs and avoiding unexpected lockouts.
The Root Problem: Context Compounding
LLMs re-read the entire session history on every turn. A 3,000-token file loaded at turn 4 of a 40-turn chat is billed 37 additional times. Over time, history re-reads can account for up to 96% of total token usage. Left unmanaged, long threads exponentially drain context and budget without adding new value.
Core Strategies to Reduce Token Consumption
- Reset Context with
/clear: The single most effective action is starting fresh between distinct tasks. Use/renameto preserve current work for later resumption via/resume, then run/clearto drop the compounded history back to zero. - Avoid Mid-Session Model and Setting Swaps: Prompt caching reduces token costs by roughly 90%, but the cache key relies on model identity, effort settings, and fast mode status. Switching from Opus to Sonnet mid-session invalidates the cache, forcing the entire conversation to reprocess at full price (up to 10x the cost). Set your model and effort levels at the start and do not alter them mid-thread.
- Filter Command Outputs: Terminal commands (such as package installations) can dump hundreds of lines of non-essential output directly into the context window. Intercept and truncate raw outputs with filter scripts before the agent ingests them.
- Prune Unused Tools and Ensure Deferral: Model Context Protocol (MCP) integrations load instruction sets into memory. Check
/contextto verify tool definitions are set to “deferred” (loading only on demand), and disable unused integrations. - Use Sub-Agents Strategically: Sub-agents do not inherently save tokens—they load their own system prompts, memory, and tools, often consuming 4x to 15x more raw tokens upfront. They only yield net savings when handling high-volume tasks whose low-level details will not be needed across subsequent turns of an extended session. Where possible, assign sub-agents to lighter models like Haiku.
- Audit Scheduled Background Tasks: Tasks operating on automated intervals resend the entire context of their assigned session. Because subscription caches expire after one hour, tasks scheduled more than an hour apart consistently miss cache hits, reprocessing the full context at peak rates.
Costly Myths and Counterproductive Habits
- Writing Shorter Prompts: Prompt length has a negligible impact on costs; vague prompts actually increase costs by triggering repetitive file searches and revisions.
- Compacting Conversations: Compaction requires a costly full-context re-read and destroys the prompt cache. To back up a few steps without breaking cache, use
/rewindinstead. - Using Screenshots and PDFs: Multi-modal inputs are expensive. A screenshot costs several thousand tokens compared to raw text, while PDFs are often processed as both text and image. Always convert source documents to plain text before analysis.
Key Takeaways and Monitoring Tools
AI providers do not optimize your token footprint by default; maintaining lean context is the user’s responsibility. Make regular use of internal inspection tools to audit your environment:
/contextto view current context breakdowns and verify deferred tool loading./usageto pinpoint specific skills, tools, or agents driving up token expenditure./costand the live terminal burn rate meter to maintain continuous visibility over spend.
Mentoring question
Looking at your current AI coding workflow, how often do you leave long conversation threads running or switch models mid-session, and what specific steps can you take today to prevent cache invalidation?
Source: https://youtube.com/watch?v=kHtOSJRUkLs&is=gAfSy9_3nHwUGQKa