data//run-costs

What one Bureau run costs

The two Bureau essay runs with complete token accounting cost 24.64 and 12.27 in API-equivalent dollars, broken down by role and token type.

figures as of Sep 6, 2026·published Sep 6, 2026·download json

This page answers one question: what does one Bureau run cost, in tokens and API-equivalent dollars, by role, across the runs that produced devweb's essays, glossary entries and topic overviews? Across the two essay runs where every leg was captured, the totals are 24.64 and 12.27 in API-equivalent US dollars, and in both of them the Conductor is the single largest line: 18.81 of the 24.64, and 6.51 of the 12.27. Every figure below is copied from the run accounting records as they were written, including the rows that recorded nothing.

Price basis

AliasPriced asInput $/MCache write $/MCache read $/MOutput $/MSourceFetched
sonnetClaude Sonnet 52.04.00.210.0https://platform.claude.com/docs/en/about-claude/pricing2026-09-06
opusClaude Opus 55.010.00.525.0https://platform.claude.com/docs/en/about-claude/pricing2026-09-06
haikuClaude Haiku 4.51.02.00.15.0https://platform.claude.com/docs/en/about-claude/pricing2026-09-06
fableClaude Fable 5.110.020.00.2550.0https://platform.claude.com/docs/en/about-claude/pricing2026-09-06
grok (OpenRouter)x-ai/grok-4.31.25n/a0.22.5https://openrouter.ai/x-ai/grok-4.32026-09-06

"API-equivalent" is carrying weight on this page, so here is what it means.

The Claude legs ran on a Claude Code subscription. They were not metered per token, and there is no invoice for them. Their dollar figures are modeled: the token counts each run recorded, multiplied by the list prices in the table above. The only legs that were actually metered are the Grok cross-model passes, which ran through OpenRouter and have a real per-call price.

Four more details that change how to read the numbers:

  • Grok token counts are estimated from byte counts at 4 bytes per token. They are marked (est) in the leg tables below, and they are not measurements.
  • Cache writes are priced at the 1-hour cache rate.
  • The model aliases (sonnet, opus, haiku, fable) are priced at the model each alias resolves to on the as-of date, 2026-09-06. Runs from June and July executed on whatever those aliases pointed at then, which may have been earlier models at different prices.
  • The token counts include cache reads, which is why a single essay run reports 27,804,254 tokens processed.

What these dollars are not

Three claims this page does not make.

Not spend. Nobody paid 24.64 for the run that produced /articles/vectors-in-the-file. That is what the recorded tokens would have cost at API list price on the table above. A subscription was the actual billing relationship, and a modeled equivalent is not an invoice.

Not comparable across providers or generations. Claude legs are priced on the Claude list, Grok legs on the OpenRouter list, and both are read as of one day. Putting a June run and a September run in the same column does not put them on the same price curve.

Not an average. Sixteen runs are in the table. Three are graded complete coverage, and two of those are essay runs. Two data points do not support an average, a per-article cost, or a trend.

The runs

RunDateWorkflowProducedCoverageLegs captured/recordedSpecialists recordedTokens processed (captured)Output tokensCache-read shareGrok callsGrok USDClaude USD equivTotal USD equiv
20260723-devweb-recallatron-improvements2026-07-23write-articlestructured-memory-graph-workerspartial10/11103,833,10615,53188%10.01454.924.93
20260723-devweb-static-nextjs-wordpress2026-07-23write-articlestatic-nextjs-vs-wordpressnone0/1000n/a10.0088n/an/a
20260905-glossary-harness-cluster2026-09-05write-glossary8 glossary-termspartial3/3035,429,14063,47994%80.018421.7721.79
20260906-glossary-memory-cluster2026-09-06write-glossary7 glossary-termspartial3/3055,082,07642,39995%70.016626.0326.05
20260906-topic-overviews2026-09-06write-topic-overview3 topic-overviewscomplete11/11850,805,09533,62693%30.009634.2534.26
20260624-delegate-gating-article2026-06-24write-articlegate-the-flow-not-the-judgmentnone0/9900n/a10.0105n/an/a
20260624-meta-pipeline-article2026-06-24write-articlethe-pipeline-that-wrote-thisnone0/9900n/a10.0113n/an/a
20260624-project-agent-article2026-06-24write-articledeny-by-default-groundingnone0/9900n/a10.0100n/an/a
20260625-growoperative-architecture-article2026-06-25write-articlegrowoperative-foaf-protocol-architecturenone0/9900n/a10.0150n/an/a
20260626-rheo-memory-v2-article2026-06-26write-articlegive-the-model-the-tools-not-the-contextnone0/9900n/a10.0098n/an/a
20260705-article-rheo-track52026-07-05write-articlevectors-in-the-filecomplete12/121127,804,254105,28395%10.010624.6324.64
20260705-article-run-metrics2026-07-05write-articleagent-run-cost-attributionpartial11/12114,466,01216,75182%10.01167.127.13
20260712-audit-convergence-article2026-07-12write-articleclean-audit-run-proves-nothingcomplete12/121112,955,27045,98095%10.011112.2612.27
20260723-bureau-shift-left-quality2026-07-23write-articleshift-left-quality-agentic-pipelinenone0/0000n/a10.0092n/an/a
20260723-devweb-wordpress-nextjs-v22026-07-23write-articlestatic-nextjs-vs-wordpresspartial1/10112,10451199%10.01090.080.09
20260723-write-article-evalpal2026-07-23write-articleevalpal-audhd-assessmentpartial2/208,539,5179,32395%10.00978.268.27

Coverage is a grade on the record, not on the run. A run grades complete when the Conductor and at least one specialist spawn are recorded, and every recorded leg carries exact token counts. Partial means only some of the recorded legs carry counts, so the totals in that row are a floor and not the run's real footprint. None means no leg carries counts at all, which is why those rows show 0 tokens and n/a dollars: there is nothing to model, and filling the gap with an estimate would make the table look better than the record is.

Two columns need reading carefully. Legs captured/recorded counts only the legs the accounting actually recorded, so a leg the run never recorded stays out of the denominator and shows as not recorded in the per-leg table below. The Delegate and the cold reviewer on runs that predate the Delegate are the common case. Specialists recorded is the number of spawn records present in the accounting, not the number of specialists that ran: a run whose accounting recorded no specialist spawns still ran them, so its totals undercount by whatever those legs spent.

The per-leg detail behind the grade:

RunSchemaCoverageConductorDelegateReviewerSpecialists captured / recordedGrok calls
20260723-devweb-recallatron-improvementsv2partialnot recordedexactnot recorded9/101
20260723-devweb-static-nextjs-wordpressv2nonenot recordedpartialnot recorded0/01
20260905-glossary-harness-clusterv2partialexactexactexact0/08
20260906-glossary-memory-clusterv2partialexactexactexact0/07
20260906-topic-overviewsv2completeexactexactexact8/83
20260624-delegate-gating-articlev1noneabsentabsentabsent0/91
20260624-meta-pipeline-articlev1noneabsentabsentabsent0/91
20260624-project-agent-articlev1noneabsentabsentabsent0/91
20260625-growoperative-architecture-articlev1noneabsentabsentabsent0/91
20260626-rheo-memory-v2-articlev1noneabsentabsentabsent0/91
20260705-article-rheo-track5v2completeexactabsentabsent11/111
20260705-article-run-metricsv2partialpartialabsentabsent11/111
20260712-audit-convergence-articlev2completeexactnot recordednot recorded11/111
20260723-bureau-shift-left-qualityv1noneabsentabsentabsent0/01
20260723-devweb-wordpress-nextjs-v2v2partialnot recordedexactnot recorded0/01
20260723-write-article-evalpalv2partialexactexactnot recorded0/01

The Schema column explains most of the empty rows. Six runs are schema v1, which recorded no per-leg token data at all: the five from June 2026, plus 20260723-bureau-shift-left-quality. Those articles exist and shipped, but the runs that made them left no accounting behind. One more row, 20260723-devweb-static-nextjs-wordpress, is schema v2 and still captured nothing usable, at 0 of the 1 leg it recorded. The remaining six v2 rows record something without clearing the bar, and they fall short in two different ways. 20260723-devweb-recallatron-improvements and 20260705-article-run-metrics each miss one of the legs they recorded, at 10 of 11 and 11 of 12. The other four captured every leg they recorded and still grade partial: none of them recorded a specialist spawn at all, which is what the 0 under Specialists recorded means. Those runs spawned specialists; the record simply does not hold them, so their totals are a floor.

Three rows are graded complete: the topic-overview cluster, which produced three topic overviews rather than one essay, and the two essay runs broken out below.

Where the cost sits

20260705-article-rheo-track5, which produced /articles/vectors-in-the-file (2026-07-05)

AgentModelInputCache writeCache readOutputUSD equivArithmetic
The Conductoropus26,512470,06923,346,12792,07018.810.027M x $5.0 + 0.470M x $10.0 + 23.346M x $0.5 + 0.092M x $25.0 = $18.81
Tallysonnet20111,8941,305,54510,4600.810.000M x $2.0 + 0.112M x $4.0 + 1.306M x $0.2 + 0.010M x $10.0 = $0.81
The Counselorsonnet1043,730143,0023170.210.000M x $2.0 + 0.044M x $4.0 + 0.143M x $0.2 + 0.000M x $10.0 = $0.21
The Scribesonnet943,169229,1643680.220.000M x $2.0 + 0.043M x $4.0 + 0.229M x $0.2 + 0.000M x $10.0 = $0.22
The Scribeopus13,11673,339283,394330.940.013M x $5.0 + 0.073M x $10.0 + 0.283M x $0.5 + 0.000M x $25.0 = $0.94
The Scribeopus13,11471,116221,730180.890.013M x $5.0 + 0.071M x $10.0 + 0.222M x $0.5 + 0.000M x $25.0 = $0.89
The Scribeopus13,11664,295273,956280.850.013M x $5.0 + 0.064M x $10.0 + 0.274M x $0.5 + 0.000M x $25.0 = $0.85
The Scribeopus13,11446,466164,261200.610.013M x $5.0 + 0.046M x $10.0 + 0.164M x $0.5 + 0.000M x $25.0 = $0.61
The Counselorsonnet1349,607324,4527670.270.000M x $2.0 + 0.050M x $4.0 + 0.324M x $0.2 + 0.001M x $10.0 = $0.27
The Counselorsonnet889,60671,2707230.380.000M x $2.0 + 0.090M x $4.0 + 0.071M x $0.2 + 0.001M x $10.0 = $0.38
The Scribesonnet528,18961,909170.130.000M x $2.0 + 0.028M x $4.0 + 0.062M x $0.2 + 0.000M x $10.0 = $0.13
The Challengeropus13,11041,05149,4834620.510.013M x $5.0 + 0.041M x $10.0 + 0.049M x $0.5 + 0.000M x $25.0 = $0.51
Grok pass 1grok-4.3~2,829 (est)n/an/a~2,822 (est)0.0106(11315/4)/1e6 x $1.25 + (11287/4)/1e6 x $2.5 = $0.0106
Total24.64Claude 24.63 + Grok 0.0106

20260712-audit-convergence-article, which produced /articles/clean-audit-run-proves-nothing (2026-07-12)

AgentModelInputCache writeCache readOutputUSD equivArithmetic
The Conductoropus87471,6489,701,85337,5196.510.001M x $5.0 + 0.072M x $10.0 + 9.702M x $0.5 + 0.038M x $25.0 = $6.51
Tallysonnet656,367116,8683830.250.000M x $2.0 + 0.056M x $4.0 + 0.117M x $0.2 + 0.000M x $10.0 = $0.25
The Counselorsonnet1033,041161,4385710.170.000M x $2.0 + 0.033M x $4.0 + 0.161M x $0.2 + 0.001M x $10.0 = $0.17
The Scribesonnet729,811129,1683650.150.000M x $2.0 + 0.030M x $4.0 + 0.129M x $0.2 + 0.000M x $10.0 = $0.15
The Scribefable13,35372,425202,4741,1711.690.013M x $10.0 + 0.072M x $20.0 + 0.202M x $0.25 + 0.001M x $50.0 = $1.69
The Scribeopus13,37157,781655,3681,8991.020.013M x $5.0 + 0.058M x $10.0 + 0.655M x $0.5 + 0.002M x $25.0 = $1.02
The Scribeopus13,35346,824160,7346590.630.013M x $5.0 + 0.047M x $10.0 + 0.161M x $0.5 + 0.001M x $25.0 = $0.63
The Scribeopus13,35748,296268,9588880.710.013M x $5.0 + 0.048M x $10.0 + 0.269M x $0.5 + 0.001M x $25.0 = $0.71
The Counselorsonnet1546,761414,0101080.270.000M x $2.0 + 0.047M x $4.0 + 0.414M x $0.2 + 0.000M x $10.0 = $0.27
The Counselorsonnet931,777162,4716550.170.000M x $2.0 + 0.032M x $4.0 + 0.162M x $0.2 + 0.001M x $10.0 = $0.17
The Scribesonnet1031,433252,8588000.180.000M x $2.0 + 0.031M x $4.0 + 0.253M x $0.2 + 0.001M x $10.0 = $0.18
The Challengeropus13,34939,42749,7859620.510.013M x $5.0 + 0.039M x $10.0 + 0.050M x $0.5 + 0.001M x $25.0 = $0.51
Grok pass 1grok-4.3~2,958 (est)n/an/a~2,958 (est)0.0111(11833/4)/1e6 x $1.25 + (11832/4)/1e6 x $2.5 = $0.0111
Total12.27Claude 12.26 + Grok 0.0111

Two things stand out, and they are the same thing seen twice.

The orchestrating leg dominates. In the first run the Conductor models out at 18.81 against a run total of 24.64; the next largest single leg is a Scribe opus pass at 0.94. In the second run the Conductor is 6.51 against a total of 12.27, and the next largest leg is the Scribe's fable pass at 1.69. Across both runs the dataset attributes 25.32 to the Conductor, 8.02 to the Scribe, 1.47 to the Counselor, 1.06 to Tally and 1.02 to the Challenger. The specialists are cheap. The agent holding the whole run in context is not.

Cache reads dominate the tokens. The Conductor's line in the first run is 23,346,127 cache-read tokens against 26,512 input and 92,070 output; in the second it is 9,701,853 against 874 and 37,519. The Cache-read share column reports 95% for both runs, and the dataset's rollup gives 0.951 across the pair. That is also why the dollar figures stay as low as they do: on the price table, an opus cache read is 0.5 per million against 5.0 for input and 25.0 for output. A run can process 27,804,254 tokens and still model out at 24.64.

What this cannot say yet

The rollup, for the two complete-coverage essay runs only:

  • complete_essay_runs: 2
  • median_usd_equiv_per_essay: 18.45
  • min_usd_equiv: 12.27
  • max_usd_equiv: 24.64
  • median_tokens_processed: 20379762
  • cache_read_share: 0.951

A median across two runs is the midpoint between them. 18.45 is not a typical run cost, because two observations cannot establish what typical is, and the two observations here are 12.27 and 24.64 on the same workflow. Read the rollup as a range that has been seen, not as a forecast.

The rest of the table is thinner still. Nine of the sixteen rows carry any token data at all, and six of those nine are graded partial, which means their totals undercount by an unknown amount. The other seven rows have nothing to report. Nothing here supports a per-article price, a cost-per-word, a comparison between workflows, or a claim that runs are getting cheaper or more expensive over time. The one comparison the data does support is within a run: which role spends, and on what.

Coverage improves as more runs complete under the current accounting schema. This page is generated from the run records and regenerated as more of them land, so the sample it reports on is whatever had been captured on the as-of date, 2026-09-06. Related reading: agent run cost attribution on how the per-leg numbers get collected, and multi-agent pipeline for what the roles in the leg tables actually do.

$ cat sources.txt