Skip to content
d3 Wiki
05.04

Token usage optimization

Faster answers, lower cost, less drift — regardless of plan.

Updated
Oct 3, 2026
On this page

Purpose

Faster answers, lower cost, less drift — regardless of plan.

The standard

  • Context hygiene: /clear when switching tasks; /compact when a session grows long; don't carry a debugging session into a new feature.

  • Scope the input: reference specific files (@src/features/panels/actions.ts) rather than "look at the codebase". Point Claude at the wiki page or ADR instead of re-explaining.

  • Plan first, then execute — a 200-token plan avoids a 20k-token wrong turn.

  • Right-size effort: low for boilerplate, high for hard problems; default otherwise.

  • Let tools do the work: ask Claude to run tests and read output rather than pasting logs. Give it grep/rg instead of dumping files.

  • CLAUDE.md is cheap context — 100 good lines beats 10 re-explanations per day.

  • Sub-agents for read-heavy research so the main context stays lean.

  • Batch small edits into one request; avoid one-line back-and-forth.

  • Plan economics: heavy daily users → Max subscription; programmatic → API with prompt caching and the smallest model that passes evals (Haiku for classification, Sonnet for most, Opus for hard reasoning).

Measure

/cost in Claude Code; API usage dashboard monthly, reviewed by the AI tooling owner.

Anti-patterns

Pasting screenshots of code instead of the file; long sessions with 50 unrelated tasks; asking for "the whole file again" after a small change.

Owner: Matt · Last reviewed: 2026-09