Add balance when needed. Funds are reserved before dispatch and settled from terminal usage.
Compare API costs before you ship.
A practical comparison of prepaid blended pricing, official split tariffs and subscription-style access. Start with your real token mix, not a headline number.
Model.sale public rates
Prices are per 1M charged tokens. The initial reservation is temporary and released after terminal usage; it is not a per-request fee. Prices remain listed when a model is temporarily unavailable.
| Model | Blended / 1M | Initial reservation | Status |
|---|---|---|---|
| claude-haiku-4-5 | $0.075 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| claude-opus-4-5 | $0.30 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| claude-opus-4-6 | $0.40 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| claude-opus-4-7 | $0.50 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| claude-opus-4-8 | $0.60 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| claude-sonnet-4-5 | $0.15 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| claude-sonnet-4-6 | $0.25 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| deepseek-v4-flash | $0.20 | $0.01 (released) | LIVE |
| deepseek-v4-pro | $0.20 | $0.01 (released) | LIVE |
| glm-5 | $0.12 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| glm-5.1 | $0.16 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| glm-5.2 | $0.20 | $0.01 (released) | LIVE |
| glm-5.3 | $0.20 | $0.01 (released) | LIVE |
| glm-5.3-flash | $0.10 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| gpt-5.4 | $0.10 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| gpt-5.4-20 | $0.10 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| gpt-5.4-mini | $0.09 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| gpt-5.5 | $0.45 | $0.01 (released) | LIVE |
| gpt-5.6-luna | $0.09 | $0.01 (released) | LIVE |
| gpt-5.6-sol | $0.45 | $0.01 (released) | LIVE |
| gpt-5.6-terra | $0.15 | $0.01 (released) | LIVE |
| gpt-6-astra | $1.125 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| gpt-image-2 | $0.10 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| kimi-k2.5 | $0.07 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| kimi-k2.6 | $0.11 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| kimi-k3 | $0.24 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| mimo-v2-omni | $0.05 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| mimo-v2-pro | $0.05 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| mimo-v2.5 | $0.01 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| mimo-v2.5-pro | $0.03 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| minimax-m2.5 | $0.02 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| minimax-m2.7 | $0.04 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| minimax-m3 | $0.04 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| qwen-3.8 | $0.25 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| qwen3.5-plus | $0.05 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| qwen3.6-plus | $0.04 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| qwen3.7-max | $0.11 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| qwen3.7-plus | $0.03 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
Input, cached input and output can have different prices and context tiers.
A temporary reservation protects the wallet, then the final charge uses only terminal tokens returned by the API.
Popular API alternatives
Model.sale is a prepaid single-service catalog. The products below are useful alternatives, but they are not identical offers: some are multi-provider routers, some are direct open-model platforms, and some bill after usage. We compare the buying model and the documented trade-offs, not model quality or a universal “cheapest” claim.
| Service | Billing | Catalog / routing | Best fit | Checked |
|---|---|---|---|---|
| OpenRouter Multi-provider aggregator | Pay-as-you-go credits; 5.5% credit purchase fee (minimum $0.80); crypto purchases have a 5% fee. | 400+ models and 70+ providers on the pay-as-you-go plan. Provider selection, fallbacks and auto-routing are available. | Teams that want breadth and provider choice in one account. Trade-off: More routing and pricing dimensions to understand; a headline per-token price is not a blended workload cost. | 2026-08-31 · source |
| Together AI Open-model inference platform | Serverless inference is pay-per-token with no provisioning cost or minimum; dedicated options use reserved capacity. | Open-weight models including DeepSeek, Qwen, Kimi, GLM and Llama families. You select the model/endpoint; serverless capacity is shared and best-effort. | Developers who need open models and want a direct, programmable API. Trade-off: You still manage model selection, split input/output math and the provider account; dedicated capacity adds a different cost model. | 2026-08-31 · source |
| Fireworks AI Inference platform | Serverless inference is per-token; dedicated deployments are billed by GPU time. New users may receive free credits. | A curated serverless catalog focused on open models, with custom and dedicated deployment paths. You choose a serverless model or run a dedicated deployment; no universal multi-provider router is promised. | Fast prototypes that may later need a dedicated open-model deployment. Trade-off: GPU-hour costs and operational choices matter once you leave serverless; compare total workload cost, not only token price. | 2026-08-31 · source |
| Groq Low-latency inference platform | Developer usage is billed in arrears with progressive thresholds; the free/developer limits and payment requirements differ. | A focused catalog of production and preview models, published with model-specific limits and token rates. You select a Groq model; the platform emphasizes predictable low-latency serving rather than cross-provider routing. | Latency-sensitive applications using the models in Groq's active catalog. Trade-off: It is not a broad aggregator, and postpaid billing is a different budget model from a prepaid wallet. | 2026-08-31 · source |
Source pages and checked dates are part of the comparison. Terms, fees, model catalogs and limits can change; verify the linked documentation before making a production decision. Model.sale does not imply affiliation with any listed service.
Illustrative workload examples
These examples compare billing shapes for the same 800K input + 200K output mix. They do not claim that similarly named models have identical quality or behavior.
| Reference family | Model.sale blended | Model.sale example | Official split example |
|---|---|---|---|
| GPT-5.6 Luna | $0.09/M | $0.09 | $0.40 · $0.20 input + $1.20 output |
| GPT-5.6 Terra | $0.15/M | $0.15 | $4.00 · $2 input + $12 output |
| GPT-5.6 Sol | $0.45/M | $0.45 | $7.20 · $4 input + $20 output |
The Model.sale column applies one blended rate to 1M charged tokens. The official column applies the published input/output rates to the stated mix. Change the mix and the result changes.
Prepaid vs subscription access
| Question | Prepaid API | Subscription product |
|---|---|---|
| How you pay | Deposit balance, then pay usage | Recurring seat or plan fee |
| Best for | Variable workloads, scripts and coding agents | Predictable seats and bundled product features |
| Cost visibility | Request-level token and ledger records | Usually plan usage limits or quotas |
| Risk to budget | Set key and daily spend limits | Recurring charge until cancelled |
How to calculate your effective cost
- Export a representative hour or day of input, cached and output tokens.
- Separate retries, failed calls and streaming disconnects.
- Apply each provider’s exact input/output/cache rules and context tier.
- Apply Model.sale blended price to the measured charged-token total; the initial reservation is released after settlement.
- Compare the resulting total, latency, availability and model capability together.
Official reference rows below are informational, dated and linked to their source. They are not a promise of parity or a benchmark.
Official published rates
USD per million tokens, standard published rates. Cached input is shown separately where the source publishes it.
| Model | Input | Cached input | Output | Source |
|---|---|---|---|---|
| OpenAI GPT-5.6 Sol | $4.00 | See source | $20.00 | Official pricing |
| OpenAI GPT-5.6 Terra | $2.00 | See source | $12.00 | Official pricing |
| OpenAI GPT-5.6 Luna | $0.20 | See source | $1.20 | Official pricing |
| Anthropic Claude Opus 4.8 | $5.00 | $0.50 | $25.00 | Official pricing |
| Anthropic Claude Sonnet 4.6 | $3.00 | $0.30 | $15.00 | Official pricing |
| Anthropic Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 | Official pricing |
Not an equivalence claim
Names, context limits and quality tiers can differ. Treat the table as a billing-shape reference.
Rates can change
Check the source date and verify official pages before committing to a workload.
Need exact math?
Read the pricing methodology for measured usage, reservations and settlement.