Transparent calculator methodology
How the Halvr savings calculator works.
The calculator is a planning model, not a benchmark or guarantee. Its fixed assumptions are listed here and are separate from the request-level receipts used for production billing.
Summary
Halvr's calculator models 30 days of traffic, output tokens equal to 35% of input tokens, a $2.5 per million tokens input rate, and a $10 per million tokens output rate. It applies independent illustrative savings rates of 12% for cache, 9% for routing, and 7% for model selection, or 28% gross before the Performance fee.
Gross assumption
28%
Output ratio
35%
Status
Illustrative, not guaranteed
Traffic and tokens
Monthly request volume equals API calls per day multiplied by 30. The entered tokens per request are treated as input tokens. Estimated output tokens are 35% of total input tokens, rounded down to a whole token.
This is a simplifying assumption. Production workloads should replace it with observed input and output distributions by feature, because chat, extraction, agent, and embedding workflows have materially different shapes.
Baseline token prices
The planning baseline uses $2.5 per million tokens input tokens and $10 per million tokens output tokens. Those are neutral calculator assumptions, not a claim about the current price of a named provider model.
Production accounting does not use these blended rates. The proxy applies the verified catalog price for the resolved model or the explicit price configured for a custom endpoint, then reconciles against actual provider-reported tokens.
Savings assumptions
The calculator applies 12% of baseline provider spend to cache savings, 9% to routing savings, and 7% to model-selection savings. The components sum to 28%; they are added once and never compounded.
A real workload may have no safe cache candidates, may already use the smallest acceptable model, or may exceed these assumptions. Halvr does not present the percentage as measured customer performance.
Performance fee
The calculator applies the current 20% Performance savings share and its $99 monthly minimum to illustrative gross savings. Optimized provider spend plus the Halvr fee becomes the modeled total with Halvr.
Net customer savings is baseline provider spend minus that total, floored at zero. The minimum means a low-volume estimate can show no net customer savings even when gross provider savings are positive.
Production evidence
Production savings receipts use request evidence: requested and resolved models, actual input and output tokens, cache state, and the catalog prices active for that decision. Cache, routing, and model-selection components remain separate so operators can inspect the counterfactual.
Teams should evaluate savings alongside task quality, retries, latency, and successful product outcomes. A lower provider bill is not a saving when the change causes more failed workflows or manual review.
Try Halvr
Route a real request through Halvr.
Keep your current SDK. Change the base URL, attach customer and feature metadata, and set a spend limit.