Skip to main content
AI

5-Hour Quota Gone in Minutes: What Subagents Taught Me

This week I burned my Claude quota in minutes – with Haiku, not Sonnet or Opus. Cause: heavy subagent use on research. Three lessons on token overhead.

Automatically translated from German · Read the original

Julian Weyer
Julian Weyer July 23, 2026 · 1 min read
AI ·AI ·Claude ·1 min read

This week I “managed” it: I burned through my 5-hour Claude quota in minutes — ouch. And with Haiku, mind you. Not with Sonnet, not with Opus, not with Fable.

The reason: heavy use of subagents during a research task. Proud of it? No! Anyone who knows me knows my opinion on token consumption and the term “reactive power”. Electrical engineers know what I mean.

Lessons learned

  1. Multi-purpose agents (like Claude Code) generate considerable token overhead.
  2. For one-off, non-repeatable tasks that may be justified. For tasks that can be standardized, it mostly just burns tokens.
  3. Local LLMs can also get you out of a jam — with clear tasks they deliver surprisingly reliably.
Takeaway

Subagents always cost something: the token overhead pays off for one-off, unstructured research tasks — for anything that can be standardized, it is mostly pure reactive power. And sometimes a local LLM with a clear assignment is plenty.