Skip to content
← Back

Lesson 6 of 6

Evaluate before/after; keep the cheaper workflow only if quality holds

Run the same tasks on the baseline and optimized workflows, report every measured category plus quality, and accept the cheaper path only when the pass criteria still pass.

Hold the tasks still

Use the same three tasks and acceptance checks for both workflows. For code tasks, start each attempt from a separate disposable copy of the same starting files; do not let the second workflow inherit the first solution. Keep your working project untouched. Use comparable fresh sessions unless session reuse is the specific change being tested.

Change one factor at a time where possible: brief, context selection, cache setup, bounds, or model. Record model/version, settings, and whether caches were cold or warm. If several factors changed together, report a workflow comparison rather than attributing the outcome to one technique.

Repeat when affordable, retain failures and retries, and disclose when you have only one pair. Offline learners use the supplied simulated pairs and state that real-world savings remain untested.

Report categories, not a single “tokens saved %”

For each task and workload, fill:

  • peak per-request input and total input across all calls (correct provider accounting)
  • output (including reasoning if reported)
  • cached read / cache write
  • tool calls, attempts, stop reason
  • cost if known (xAI ticks, dashboard export, or unknown)
  • quality: pass / fail / not run, with the command or inspection

A drop in billed input that is only cache reads is a discount, not token reduction. A drop in occupancy with a failed test is not a win. Published cache and batch discounts are pricing rules, not your task-level result. Write what you measured. If cost is unknown, say unknown — do not multiply stale dollar rates from a blog.

Quality gates decide, not the cheaper column

Accept the optimized workflow for a task only if:

  1. Every pass criterion still passes (tests, or the written expected answer).
  2. You did not remove safety checks, secret rules, or tests to get the number.
  3. Stop reasons are legitimate (success or a documented bound), not “gave up and skipped asserts.”

If quality fails, keep the baseline for that task and say why. Mixed outcomes are honest: maybe thin context helped the search task and hurt the schema task. The course does not certify savings, issue credits, or promise cheaper production bills.

Write the report others can replay

A useful report lets someone else see the tasks, the two workflows, the numbers, and the keep/reject calls without trusting a slogan. Create token-lab/REPORT.md with provider and model names only, the pass criteria, a table of occupancy / output / cache categories / tools / attempts / cost-if-known, quality command output, what you changed, and a decision per task. State limitations: cache hits are not assured, list prices change, and this file is not a certificate.

# Token-usage report
Provider/tool: …  Models: … (names only)
## Tasks and pass criteria
## Baseline vs optimized (table of categories)
## Quality results
## What changed in the workflow
## Decision per task: keep optimized / keep baseline / mixed
## Limitations: not certified; rates change; cache hits not guaranteed

If you choose to save a submission, share an HTTPS URL containing only a sanitised report you have permission to publish. Keep proprietary code and raw private logs out; a small report is enough. You can study and practise without publishing or signing in. Saving the course form is self-reported pending review — not verification of the report or a promise of instructor review.

Sources

How did this lesson go? Give feedback →