FIELD GUIDE / MODEL ROUTING · SEPTEMBER 2026
Spend on the
decisions that matter.
A starting guide to cost-conscious distributed-system engineering with Luna, Terra, Sol, and Astra.
THE OPTIMIZATION TARGETTotal cost per
accepted result.
01 / CHOOSE THE SCOPE
Match the model to the work.
These are hypotheses to test on your own engineering tasks, not capability guarantees.
Swipe or scroll the table to see examples and checks →
| Start with | Good starting scope | Example assignment | Acceptance check |
|---|---|---|---|
| LunaMechanical work | Narrow, easily verified enumeration and summaries of verified findings. | List timeout configuration and every call site that uses it. | Complete coverage, file references, and no unsupported conclusions. |
| TerraLocal implementation | A clear contract, a bounded change, and a known expected result. | Implement specified retry and timeout behavior; add a known-bug regression test. | The regression fails before the fix and passes after; the contract holds. |
| SolAcross components | Implementation or investigation spanning several parts of a system. | Trace duplicate message processing and implement an established idempotency design. | Evidence connects the components; replay and failure checks cover the design. |
| AstraArchitecture & judgment | Ambiguity, consistency guarantees, transaction boundaries, ordering, recovery, and critical review. | Decide where the transaction ends and how recovery preserves the invariant. | Explicit assumptions, tradeoffs, failure scenarios, and defensible guarantees. |
02 / KEEP THE LOOP SHORT
Delegate a contract.
Review the evidence.
Use meaningful checkpoints: after investigation, before a design becomes implementation, and before accepting the result.
Worker + coordinator + retries
+ review & rework
Measure comparable total spend. Track human review time and latency alongside it.
Establish the invariants.
State required behavior before choosing a model. What must remain true under duplicate delivery, timeout, partial failure, and recovery?
Delegate one bounded outcome.
Provide relevant context, scope, constraints, and acceptance checks. Make the worker’s authority to change the design explicit.
Ask for a short, evidenced report.
Require findings or changes, checks performed, file references, and unresolved uncertainty. Concision should preserve the evidence.
Correct once, then reassess.
Give one targeted correction. Escalate repeated failure or new design decisions; do not keep retrying a task whose scope has changed.
Accept against the contract.
Review at the planned checkpoints. Count a result as accepted only when its checks and required behavior are satisfied.
03 / A SMALL EXPERIMENT
Equal scores. Different costs.
Ten simple trivia questions per model. Three fresh agents each. Medium reasoning throughout.
All four models
All answers correct
What the run demonstrated
Luna had the lowest API-equivalent estimate in this trial. No corrections or retries were needed. Simple trivia did not expose a quality advantage for the larger models.
This is not an engineering benchmark. The routing policy above is a proposed starting point, not a finding of this experiment.
Swipe or scroll the table to see every column →
| Model | Total tokens | API-equivalent cost | Relative to Luna | Score |
|---|---|---|---|---|
| Luna | 148,345 | $0.00864288 | 1.00× | 10/10 |
| Terra | 107,072 | $0.055758 | 6.45× | 10/10 |
| Sol | 107,067 | $0.1117368 | 12.93× | 10/10 |
| Astra | 94,301 | $0.140242 | 16.23× | 10/10 |
Matched prompts ≠ matched context
Corresponding agents received identical assigned prompts, with no conversation-history fork and no tools. Full input sizes and cache coverage still differed; their cause was not isolated. Cache coverage especially favored Astra. Ratios are specific to these runs.
Estimates ≠ subscription deductions
Standard short-context API rates were applied to measured counters. Coordinator and voice usage are excluded. These figures are not actual Codex charges or subscription allowance deductions.
Inspect counters, rates, and assumptions
| Model | Uncached input | Cached input | Output | Reasoning subset |
|---|---|---|---|---|
| Luna | 29,206 | 118,784 | 355 | 249 |
| Terra | 18,693 | 88,320 | 59 | 0 |
| Sol | 18,815 | 88,192 | 60 | 0 |
| Astra | 4,767 | 89,472 | 62 | 0 |
Medium reasoning is a setting, not a fixed token budget. Cache-write counts were zero. Each request was below the long-context threshold. No regional surcharge or fast-mode premium is assumed.
| Model / official source | Uncached input | Cached input | Output |
|---|---|---|---|
| Luna ↗ | $0.20 | $0.02 | $1.20 |
| Terra ↗ | $2.00 | $0.20 | $12.00 |
| Sol ↗ | $4.00 | $0.40 | $20.00 |
| Astra ↗ | $10.00 | $1.00 | $50.00 |
Estimate = (uncached input × input rate + cached input × cache rate + output × output rate) / 1,000,000. Rates are the report’s dated assumptions, not a live price feed.
Hypothetically removing all cache discounts yields Luna $0.030024, Terra $0.214734, Sol $0.429228, and Astra $0.945490. This is a sensitivity calculation, not a second measurement; input sizes still differ.
The ten questions were divided 4 / 3 / 3 across each model’s agents. Each agent replied once. Earlier exploratory runs are excluded. Download measured counters (JSON); private agent identifiers are omitted.
Read the full experiment.
Two-page report with method, counters, rates, and limitations.
04 / VALIDATE ON REAL WORK
Make your own acceptance rate
the next data point.
Repeat representative engineering tasks with controlled context and cache conditions. Record accepted outcomes, retries, review effort, latency, and coordinator overhead. Update the routing policy when that evidence changes.