Finest vs pinning one model
Last updated 2026-08-19
Pinning one strong model for all traffic is simple, predictable, and the right call more often than vendors admit: one integration, one behavior, no routing to reason about. Its cost is structural: every easy request pays the hardest request's price, forever. Finest keeps the part of pinning that matters, your chosen model as the verbatim default, and removes the flat tax: where sealed evidence proves a cheaper configuration holds the quality bar for a task class, it serves that instead and charges 25% of the difference it just proved. Unproven traffic stays exactly as pinned.
| At a glance | Finest | Pinning one model |
|---|---|---|
| Simplicity | One base URL swap, then automatic | Unbeatable: one model ID in code |
| Cost profile | Each task class pays its proven price | Every request pays the frontier price |
| Quality risk | Bounded by published bars and escalation | None added; the pin is the ceiling too |
| Fee | No model markup. 25% of proven savings per request. No saving, no fee. | None, and no savings either |
| Visibility | A receipt per request: model served, evidence, saving. | A provider invoice |
| Reversibility | FINEST_DISABLE=1, one line | Nothing to reverse |
Choose Finest when
- Real spend, mixed workload: extraction and classification riding at frontier prices next to genuinely hard work.
- You want savings that arrive with receipts you can put in front of finance or a customer.
- You want the experiment priced at zero: no proven saving on your traffic, no fee.
Choose Pinning one model when
- Spend is too small for optimization to matter yet. Pin the best model and build.
- The workload is uniformly frontier-hard, so there is little differential to capture.
- A dependency freeze or compliance posture forbids any vendor in the request path, even one with a one-line exit.
What the pin actually costs
Frontier and efficient models are priced far apart, and published prices move fast enough that the exact multiple is a live number rather than a slogan. The strip at the bottom of this page reads current published rates from the same registry Finest prices with.
A pinned model converts that spread into a flat tax. The requests that needed the frontier get it, and so do the requests that did not: the metadata extraction, the classification, the reformatting that makes up most production traffic. Nobody chose that spend. It is the default nobody got around to examining, which is why it is usually the largest saving available. The Request Compiler, Finest's published method, reports what examining it was worth on one real 82-document pipeline: 48% lower cost and 26% lower latency with zero added hallucinations on about 90% of documents, with the genomic-report class that failed its bar honestly kept on the frontier model.
Finest treats your pin as the contract
The reason pinning feels safe is that it is a promise: this model, exactly. Finest keeps the promise. The requested model is served verbatim unless the full request configuration matches a sealed, published evidence bar authorizing a cheaper serve for that task class, and a cheaper arm that refuses or fails a validator escalates back to your pin. The pin stops being a ceiling on your economics while remaining the floor under your quality.
Questions people ask
- When is pinning one model the right choice?
- Small spend, uniformly hard workloads, or a hard requirement of zero vendors in the request path. Pinning is a legitimate strategy, and the honest baseline every router should be measured against.
- Does Finest ever override an explicit pin without evidence?
- No. Unproven traffic serves the requested model verbatim, at the host's rate, with no markup and no fee.
- How do I find out what my pin is costing me?
- Route traffic through Finest and read the receipts: each one carries the counterfactual cost pinned to your model next to what actually served. If the difference is zero, Finest is free.
In 2 minutes, start cutting your API spend without sacrificing quality. Free if you don’t save money.
No model markup. You pay the host’s rate. 25% of what it proves it saved on a request. No saving, no fee.