Measured · AI usage
What AI-assisted engineering actually costs
Every figure on this page is derived from the provider's own usage exports for2026-07-31 to 2026-09-16 — 18,966 requests, 2.6B tokens,$44.21. Nothing is extrapolated and no usage is excluded: it is the practice's complete AI consumption for the period.
What the work was
Concrete deliverable counts from version control across the same period. No project names and no per-repository split — the work is characterised by kind, not by client or product.
Work profile — software engineering across independent codebases: a desktop application built on a Rust core with a GTK front end, web front ends, CLI tooling and developer libraries, alongside research, design and business operations.
Deliverables against spend
Cost tracks shipped work rather than idle consumption — both series measured over the same 48 days.
What it cost
Provider-reported spend. Cost is always taken from the rate actually charged on each line of the export, never recomputed from a price list — the model line was relabelled on 10 September 2026 (V4 Flash → V4.1 Flash) and prices moved during the period.
Daily spend
$44.21 over 48 days — $0.921 per day, every day of the period used.
Per model line
One series per line, continuous across the 10 September relabel.
| Model line | Tokens | Generated | Requests | Cost | Cache reuse | Tokens / $ | Generated / $ |
|---|---|---|---|---|---|---|---|
| Flash tier (V4 → V4.1) | 2.5B | 18,279,006 | 17,647 | $34.83 | 99.0% | 70.5M | 524,851 |
| Pro tier (V4 Pro) | 167M | 1,373,052 | 1,319 | $9.38 | 98.3% | 17.8M | 146,368 |
What the method bought
The practice's AI work runs on a convergent working profile: context that is stable, revised in place and reused, rather than rebuilt per request. The provider charges a small fraction of the input rate for context it can serve from cache, so a stable context is what makes long-context work affordable. These three charts are that effect, measured.
Cache reuse
Share of all input tokens served from the prompt cache — 98.98% over the period, and it rises as the profile converges.
What the tokens are
Re-read context (cache hits) against freshly sent input and generated text — three orders of magnitude apart on a log scale. 98.24% of all tokens moved are re-read context; 20M tokens were actually generated.
The same workload without cache reuse
Identical token volume with every input token charged at the observed cache-miss rate instead: 9.6× to 28.2× what was actually paid.
How to read these numbers
- Cost per commit ($0.046). The denominator is non-merge commits in the same period, counted from version control. It is unweighted — a documentation commit counts like a feature commit — so it is a headline comparison, not a productivity metric.
- Tokens per dollar is a cache figure. 98.98% of input is served from cache because the context is reused, so the blended59.3M tokens per dollar measures context reuse as much as work. The cache-independent figure is 444,538 generated tokens per dollar.
- Nothing is filtered. All requests, models and clients in the period are included, including research, planning and failed attempts. A benchmark that removes the unproductive work is a brochure.
- Source. Provider usage exports, ingested row by row; spend is the sum of the rates actually charged, cross-checked against the provider's own cost column to the cent. Figures regenerated 2026-09-16.