Lesson 5 of 6
Bound loops, tools, models, and batch jobs
Cap retries and expensive tool calls, pick a model that matches the task, and use asynchronous batch APIs only for independent, latency-tolerant work — without removing checks.
The loop is often the real token burner
One oversized prompt is visible. Twenty identical greps, three “fix” rounds that change nothing, and a model that re-reads the tree each time are worse. Before you start: step cap, allowed commands, stop rules (success, repeated failure, budget, ambiguity). OpenCode’s agent.<name>.steps is an iteration cap, not a dollar kill-switch; opencode stats reports usage after the fact. Other harnesses need the same limits in config and in the brief.
Stop conditions stay. Do not raise the cap because the agent looks busy. Do not delete tests so the loop goes green. Stopping avoids further uncontrolled work; report the failed check and usage already incurred. It does not prove the overall task was cheaper.
Make tool calls scarce and specific
Each tool round adds output (the call) and input (the result) to the next request. Wide find, full-file reads, and web fetches of pages you will not cite are occupancy pumps. Prefer: one search with a tight pattern, a line range, a baseline test before editing, then targeted tests after relevant changes. Server-side tools (web search, code execution) may add per-invocation charges on top of tokens — xAI documents token cost plus tool-invocation cost; Anthropic bills web search per search plus tokens. Log tool_calls as its own column.
If the same tool and input repeat without a changed file or new external state, pause and inspect the reason. A documented status poll or a test after an edit is not automatically a loop. Caching tool definitions (stable schemas at the front of the prefix) is reasonable; caching a unique 200-line stack trace is not.
Match model difficulty; do not assume cheaper is cheaper
A smaller or faster model can reduce price per token and sometimes output length. It can also increase retries and tool thrash, which dominates the bill. Providers also tokenize differently: Anthropic notes newer tokenizers may emit more tokens for the same text — recount on the model you will use.
Rule of thumb, not a ranking: use the small model for classification, extraction, and “rewrite this comment”; keep the stronger model for multi-file reasoning you already bounded. Compare task totals (all attempts, all tools), not one successful completion. Look up current per-million-token rates on the provider pricing page; this course does not freeze a dollar figure.
Batch independent requests; do not batch a live agent
Batch APIs queue independent requests for later processing. They can offer discounted token rates and separate capacity, but the supported models, prices, completion windows, and expiration behaviour vary. Check the selected provider's current Batch documentation before submitting. A queued job can expire or fail; save item IDs, match each result to its input, and count any retried items in the total.
Good candidates include independent classifications, offline summaries, and evaluations. An interactive plan–edit–test loop needs each result before deciding the next action, so it is not a ready-made batch. Jobs writing the same files also need coordination.
Batch discounts are not fewer tokens per item. Parallel synchronous requests are another mechanism and are not automatically billed at batch rates. The exercise does not require submitting a batch or buying credits.