Model selection

The cheapest API isn't the cheapest decision.

Picking a model on its pricing page optimizes the one number that lies. TokenWake ranks every configuration by true cost per delivered task — API spend plus modeled failure exposure — and recommends the one that actually costs least. It's rarely the cheapest on the sticker. Here are the results, then how we get there.

The gap

The pricing page is not the cost.

What the dashboard shows

Vendor pricing pages and usage dashboards report API spend — the sticker price, and only after the tokens are already gone.

Why cheap gets expensive

True cost per delivered task includes retries, abandonment, silent failures, and the human review and rework nobody puts on the invoice.

The hard dollar facts

Cheapest API price ≠ cheapest decision.

A sweep of 144 configurations, each ranked on true cost per delivered task. Two of them, side by side:

Winning config · rank #1

API spend

$313

True cost / delivered task

$5,801

Cheapest by API · rank #144

API spend

$5.97

True cost / delivered task

$34,654

On the pricing page, the #144 option looks about 50× cheaper. In true cost, it's 6× more expensive — the winning configuration cost a fraction of the cheapest-looking one.

How we get there

Declare → Sweep → Rank → Prune → Report

A recommendation isn't one model — it's a configuration: each step of your workflow gets its own model assignment. We test every combination of model-per-step, because the best model for one step is rarely the best for the next.

Stage 1

Declare

You declare the workflow: its steps, volumes, retry policy, budget gate, and the cost of a silent failure. Nothing is inferred from a black box. Every assumption is visible — change one, re-run, and you get a new answer.

Stage 2

Sweep

Every model is tried in every position, and every combination across the workflow is evaluated — so each step ends up with its own model assignment. A two-step draft-then-validate flow alone produces 144 combinations; a longer workflow, far more. It's exhaustive by construction, not a hand-picked shortlist — which is how a cross-vendor pairing no single vendor would recommend can surface as the winner.

Stage 3

Rank

Each configuration is ranked by true cost per delivered task — API spend plus modeled failure exposure: retries, abandonment, silent failures, and human review and rework.

Stage 4

Prune

The results are pruned using standard statistical methods — Monte Carlo ensemble, Pareto ranking, and sensitivity screens — so leadership decides among the defensible few, not 144 options.

Stage 5

Report

What ships is a confidence-rated shortlist of top configurations, each with its flip points — the thresholds at which the recommendation would change. The answer tightens further once your own operating data replaces the benchmark priors.

Statistical methods

Named, so you can check our work.

Monte Carlo ensemble

Simulation runs the workflow thousands of times across the range of each declared assumption, producing a distribution of outcomes rather than a single point estimate.

Pareto ranking

Configurations are ranked by dominance across cost and reliability at once, isolating the non-dominated frontier from the options that are simply worse.

Sensitivity screens

Each assumption is shaken ±20% to find the flip points — the thresholds at which the recommendation would change. Fragile winners are flagged as fragile.

Determinism

An auditor reruns it and gets the same numbers.

Give it the same inputs and it returns the same result, every time — down to the dollar. A forecast isn't a one-off opinion you have to take on faith; it's a computation your own governance team can replay and confirm. That's the model-governance hook: reproducibility, not persuasion.

Send us one workflow

We'll find the configuration that actually costs least.

$15,000 introductory offer · one workflow.