Lesson 1 of 6
Measure a baseline before you optimise
Record input, output, and cached tokens plus tool calls and retries for real work, and keep context-window occupancy distinct from billed usage.
Four meters that are not the same number
Track four different things: context size in one request, usage charged across requests, rate limits on throughput, and spend controls on cost. Cached input still takes context space. Repeated requests can reuse that space while adding more billed usage.
Do not assume a dashboard budget stops traffic. A spend alert only notifies; an enforced limit can reject requests. Availability and enforcement delays depend on the provider. OpenAI documents both controls and warns that a hard limit can slightly overshoot. Check the actual setting, keep a small exercise allowance, and stop manually if usage is unclear.
The course is free, but optional provider calls may incur charges. Use existing redacted logs or the simulated fixtures below without opening an API account. Label simulated work honestly; it proves your analysis, not personal savings.
Read the usage object your tool already returns
You do not need paid extra software. After a run, copy usage from the API response, the provider Usage dashboard, or your coding agent’s own report (for example opencode stats if you already use OpenCode). Field names differ:
- OpenAI / xAI Chat Completions:
prompt_tokens,completion_tokens,prompt_tokens_details.cached_tokens. Cached tokens are a subset ofprompt_tokens, not an extra pile. - OpenAI / xAI Responses:
input_tokens,output_tokens,input_tokens_details.cached_tokens. Some OpenAI model families also reportcache_write_tokens. - Anthropic Messages:
input_tokensis only the tokens after the last cache breakpoint. Total input iscache_read_input_tokens + cache_creation_input_tokens + input_tokens. - xAI also returns
cost_in_usd_ticks(divide by 10¹⁰ for USD) after discounts, including cache.
Reasoning tokens, where present, are billed as output even when they are not visible answer text. For OpenAI-compatible Chat Completions streaming, request stream_options.include_usage and retain the final usage chunk. Responses and other SDKs have different completion events; follow the endpoint documentation, not a copied Chat parameter. An interrupted stream may leave usage incomplete.
Log occupancy, bill categories, tools, and retries — never keys
Create token-lab/usage-log.jsonl. One JSON object per attempt, not per wish. Redact API keys, .env values, and auth files. Keep model names, not secrets.
{"run_id":"R1","task_id":"T1","workload":"baseline","attempt":1,"provider":"","model":"","input_tokens":null,"output_tokens":null,"cached_read_tokens":null,"cache_write_tokens":null,"reasoning_tokens":null,"total_tokens_reported":null,"tool_calls":0,"stop_reason":"","quality":"unchecked","cost_if_known":"unknown","notes":""}
Also count tool calls (grep, read, web, tests) and attempts (the same job retried). A cheap first call plus three failed fix rounds is not a cheap job. If a field is missing from your tool, write not reported — do not invent it. Context occupancy is the tokens the model had to hold. Billing categories are how those tokens were priced. Do not subtract cached tokens from occupancy.
Offline option — supplied simulated fixtures, not measured savings. Each row represents one fictional API request with no cache use. Save these rows separately as fixtures.jsonl; never relabel them as your provider usage.
{"task":"T1","variant":"baseline","input":1200,"output":300,"tool_calls":3,"quality":"pass"}
{"task":"T1","variant":"optimized","input":800,"output":250,"tool_calls":3,"quality":"pass"}
{"task":"T2","variant":"baseline","input":2000,"output":400,"tool_calls":4,"quality":"pass"}
{"task":"T2","variant":"optimized","input":1000,"output":300,"tool_calls":2,"quality":"fail"}
{"task":"T3","variant":"baseline","input":1500,"output":250,"tool_calls":2,"quality":"pass"}
{"task":"T3","variant":"optimized","input":1400,"output":350,"tool_calls":4,"quality":"pass"}
T1 adds empty-title validation; T2 enforces a 160-character limit but the optimized version wrongly accepts 161; T3 extracts three facts from one document. Use these cases to practise a comparison without making network calls. Costs and timing are unknown. Keep your written task criteria and fixture analysis as evidence.
Normalise fields without mixing accounting systems
Choose the accounting format explicitly; do not guess it from missing fields. On OpenAI-compatible usage, cache categories sit inside total input. Anthropic reports separate input, cache-read, and cache-write counts. A missing count is unknown, not zero.
Save this as token-lab/normalize-usage.mjs. Pass the saved usage object, not the entire response. It reads no secrets and makes no network calls.
const count = v => Number.isInteger(v) && v >= 0 ? v : null;
export function normalize(u, format) {
if (!['anthropic', 'openai-compatible'].includes(format)) {
throw new Error('Choose the documented usage format');
}
const anth = format === 'anthropic';
const cached = count(anth ? u.cache_read_input_tokens
: (u.input_tokens_details?.cached_tokens ?? u.prompt_tokens_details?.cached_tokens));
const cacheWrite = count(anth ? u.cache_creation_input_tokens
: (u.input_tokens_details?.cache_write_tokens ?? u.prompt_tokens_details?.cache_write_tokens));
const input = count(u.input_tokens ?? u.prompt_tokens);
const occupancyIn = anth
? ([input, cached, cacheWrite].every(v => v !== null) ? input + cached + cacheWrite : null)
: input;
return { occupancyIn, output: count(u.output_tokens ?? u.completion_tokens),
cached, cacheWrite, accounting: format };
}
Check it: normalize({input_tokens:8000,input_tokens_details:{cached_tokens:7200}},'openai-compatible') keeps input at 8000 and output unknown. normalize({},'anthropic') must return unknown counts. Explicit zeros in complete provider responses remain zero.
For a multi-call task, record peak per-request input separately from sum of input across calls. The sum measures repeated usage; it is not one request's context-window occupancy.