This week I “managed” it: I burned through my 5-hour Claude quota in minutes — ouch. And with Haiku, mind you. Not with Sonnet, not with Opus, not with Fable.
The reason: heavy use of subagents during a research task. Proud of it? No! Anyone who knows me knows my opinion on token consumption and the term “reactive power”. Electrical engineers know what I mean.
Lessons learned
- Multi-purpose agents (like Claude Code) generate considerable token overhead.
- For one-off, non-repeatable tasks that may be justified. For tasks that can be standardized, it mostly just burns tokens.
- Local LLMs can also get you out of a jam — with clear tasks they deliver surprisingly reliably.
Subagents always cost something: the token overhead pays off for one-off, unstructured research tasks — for anything that can be standardized, it is mostly pure reactive power. And sometimes a local LLM with a clear assignment is plenty.