Lesson 2 of 6
Write a compact brief and bound the output
Shorten the task brief and cap output without dropping acceptance criteria, safety rules, or the checks that prove the work.
Short is not the same as clear
A rambling prompt wastes input tokens. A vague short prompt often wastes more: the model guesses, you correct it, and retries multiply input, output, and tool calls. Compact means every sentence earns its place — outcome, constraints, allowed files, and how you will check the result — not that the text is twelve words.
Keep: observable acceptance criteria, out-of-scope lines, secret-handling rules, and the test or inspection you will run. Cut: repeated pep talks, pasted stack traces you already summarised, and “also maybe also” extras. If two readers would disagree about done, the brief is too short in the wrong place. Compare short vs clear, not short at any price.
Bound what the model is allowed to generate
Usage fields (completion_tokens or output_tokens) report generation; request parameters (max_completion_tokens, max_output_tokens, or max_tokens, depending on API) limit it. Those are different fields. Reasoning may consume the generation allowance even when it is not visible in the final answer. Caps limit generation; they are not prepaid blocks. You are charged for tokens actually produced.
Ask for the shape you need: a diff, a table, or “three bullets, then stop.” Forbid a full-file rewrite when a patch will do. Do not set a tiny cap that truncates JSON or a test plan — that creates another round. OpenAI documents that max_output_tokens limits all generated tokens, including non-visible formatting, so leave headroom when you need a complete visible answer. A prose request such as “three bullets” is a preference, not an enforced API limit. Check the stop reason and validate the output before treating it as complete.
A compact brief still names the contract
Use this shape. Adapt names; keep every row that names a check or a forbid.
# Brief — <task id>
Outcome: one sentence.
Allowed files: …
Do not change: tests that already pass, secrets, deploy, git push.
Acceptance:
- Given … when … then …
Stop when: checks pass, or the same check fails twice, or the step cap hits.
Output: patch or file list, plus the command you ran. No essay.
The “do not change” line is load-bearing. Deleting “run the tests” to save a few input tokens is not an optimisation; it removes the stop condition and usually increases later spend when you debug by hand. A bounded agent loop already treats test weakening as a failed run. The same rule applies when the motive is cost: keep acceptance criteria, secret-handling, and the named check. Cut repetition and open-ended essays, not the contract.
Compare two briefs on the same task
Take one baseline task from Lesson 1. Write Brief A (your old chatty prompt) and Brief B (the template above) without removing criteria. Run B once if you can; if you cannot re-run yet, estimate only the input size of A vs B using your tool’s tokenizer, Anthropic POST /v1/messages/count_tokens, or OpenAI POST /v1/responses/input_tokens. Those counters estimate input; they do not predict output.
Example / simulated: Brief A ~1 900 input tokens of instructions; Brief B ~420; same four acceptance lines. Do not treat that pair as your measurement. Record whether B still lists the quality check. If B is shorter because it dropped “reject empty titles” or “do not skip tests,” it failed this lesson even if the token count fell.