Independent price check · 2026-08-31

Compare API costs before you ship.

A practical comparison of prepaid blended pricing, official split tariffs and subscription-style access. Start with your real token mix, not a headline number.

Calculate your workload

Model.sale public rates

Prices are per 1M charged tokens. The initial reservation is temporary and released after terminal usage; it is not a per-request fee. Prices remain listed when a model is temporarily unavailable.

ModelBlended / 1MInitial reservationStatus
claude-haiku-4-5$0.075$0.01 (released)TEMPORARILY UNAVAILABLE
claude-opus-4-5$0.30$0.01 (released)TEMPORARILY UNAVAILABLE
claude-opus-4-6$0.40$0.01 (released)TEMPORARILY UNAVAILABLE
claude-opus-4-7$0.50$0.01 (released)TEMPORARILY UNAVAILABLE
claude-opus-4-8$0.60$0.01 (released)TEMPORARILY UNAVAILABLE
claude-sonnet-4-5$0.15$0.01 (released)TEMPORARILY UNAVAILABLE
claude-sonnet-4-6$0.25$0.01 (released)TEMPORARILY UNAVAILABLE
deepseek-v4-flash$0.20$0.01 (released)LIVE
deepseek-v4-pro$0.20$0.01 (released)LIVE
glm-5$0.12$0.01 (released)TEMPORARILY UNAVAILABLE
glm-5.1$0.16$0.01 (released)TEMPORARILY UNAVAILABLE
glm-5.2$0.20$0.01 (released)LIVE
glm-5.3$0.20$0.01 (released)LIVE
glm-5.3-flash$0.10$0.01 (released)TEMPORARILY UNAVAILABLE
gpt-5.4$0.10$0.01 (released)TEMPORARILY UNAVAILABLE
gpt-5.4-20$0.10$0.01 (released)TEMPORARILY UNAVAILABLE
gpt-5.4-mini$0.09$0.01 (released)TEMPORARILY UNAVAILABLE
gpt-5.5$0.45$0.01 (released)LIVE
gpt-5.6-luna$0.09$0.01 (released)LIVE
gpt-5.6-sol$0.45$0.01 (released)LIVE
gpt-5.6-terra$0.15$0.01 (released)LIVE
gpt-6-astra$1.125$0.01 (released)TEMPORARILY UNAVAILABLE
gpt-image-2$0.10$0.01 (released)TEMPORARILY UNAVAILABLE
kimi-k2.5$0.07$0.01 (released)TEMPORARILY UNAVAILABLE
kimi-k2.6$0.11$0.01 (released)TEMPORARILY UNAVAILABLE
kimi-k3$0.24$0.01 (released)TEMPORARILY UNAVAILABLE
mimo-v2-omni$0.05$0.01 (released)TEMPORARILY UNAVAILABLE
mimo-v2-pro$0.05$0.01 (released)TEMPORARILY UNAVAILABLE
mimo-v2.5$0.01$0.01 (released)TEMPORARILY UNAVAILABLE
mimo-v2.5-pro$0.03$0.01 (released)TEMPORARILY UNAVAILABLE
minimax-m2.5$0.02$0.01 (released)TEMPORARILY UNAVAILABLE
minimax-m2.7$0.04$0.01 (released)TEMPORARILY UNAVAILABLE
minimax-m3$0.04$0.01 (released)TEMPORARILY UNAVAILABLE
qwen-3.8$0.25$0.01 (released)TEMPORARILY UNAVAILABLE
qwen3.5-plus$0.05$0.01 (released)TEMPORARILY UNAVAILABLE
qwen3.6-plus$0.04$0.01 (released)TEMPORARILY UNAVAILABLE
qwen3.7-max$0.11$0.01 (released)TEMPORARILY UNAVAILABLE
qwen3.7-plus$0.03$0.01 (released)TEMPORARILY UNAVAILABLE
Prepaid
Pay as you go

Add balance when needed. Funds are reserved before dispatch and settled from terminal usage.

Official APIs
Usually split rates

Input, cached input and output can have different prices and context tiers.

Exact usage
Measured settlement

A temporary reservation protects the wallet, then the final charge uses only terminal tokens returned by the API.

Popular API alternatives

Model.sale is a prepaid single-service catalog. The products below are useful alternatives, but they are not identical offers: some are multi-provider routers, some are direct open-model platforms, and some bill after usage. We compare the buying model and the documented trade-offs, not model quality or a universal “cheapest” claim.

ServiceBillingCatalog / routingBest fitChecked
OpenRouter
Multi-provider aggregator
Pay-as-you-go credits; 5.5% credit purchase fee (minimum $0.80); crypto purchases have a 5% fee.400+ models and 70+ providers on the pay-as-you-go plan. Provider selection, fallbacks and auto-routing are available.Teams that want breadth and provider choice in one account.
Trade-off: More routing and pricing dimensions to understand; a headline per-token price is not a blended workload cost.
2026-08-31 · source
Together AI
Open-model inference platform
Serverless inference is pay-per-token with no provisioning cost or minimum; dedicated options use reserved capacity.Open-weight models including DeepSeek, Qwen, Kimi, GLM and Llama families. You select the model/endpoint; serverless capacity is shared and best-effort.Developers who need open models and want a direct, programmable API.
Trade-off: You still manage model selection, split input/output math and the provider account; dedicated capacity adds a different cost model.
2026-08-31 · source
Fireworks AI
Inference platform
Serverless inference is per-token; dedicated deployments are billed by GPU time. New users may receive free credits.A curated serverless catalog focused on open models, with custom and dedicated deployment paths. You choose a serverless model or run a dedicated deployment; no universal multi-provider router is promised.Fast prototypes that may later need a dedicated open-model deployment.
Trade-off: GPU-hour costs and operational choices matter once you leave serverless; compare total workload cost, not only token price.
2026-08-31 · source
Groq
Low-latency inference platform
Developer usage is billed in arrears with progressive thresholds; the free/developer limits and payment requirements differ.A focused catalog of production and preview models, published with model-specific limits and token rates. You select a Groq model; the platform emphasizes predictable low-latency serving rather than cross-provider routing.Latency-sensitive applications using the models in Groq's active catalog.
Trade-off: It is not a broad aggregator, and postpaid billing is a different budget model from a prepaid wallet.
2026-08-31 · source

Source pages and checked dates are part of the comparison. Terms, fees, model catalogs and limits can change; verify the linked documentation before making a production decision. Model.sale does not imply affiliation with any listed service.

Illustrative workload examples

These examples compare billing shapes for the same 800K input + 200K output mix. They do not claim that similarly named models have identical quality or behavior.

Reference familyModel.sale blendedModel.sale exampleOfficial split example
GPT-5.6 Luna$0.09/M$0.09$0.40 · $0.20 input + $1.20 output
GPT-5.6 Terra$0.15/M$0.15$4.00 · $2 input + $12 output
GPT-5.6 Sol$0.45/M$0.45$7.20 · $4 input + $20 output

The Model.sale column applies one blended rate to 1M charged tokens. The official column applies the published input/output rates to the stated mix. Change the mix and the result changes.

Prepaid vs subscription access

QuestionPrepaid APISubscription product
How you payDeposit balance, then pay usageRecurring seat or plan fee
Best forVariable workloads, scripts and coding agentsPredictable seats and bundled product features
Cost visibilityRequest-level token and ledger recordsUsually plan usage limits or quotas
Risk to budgetSet key and daily spend limitsRecurring charge until cancelled

How to calculate your effective cost

  1. Export a representative hour or day of input, cached and output tokens.
  2. Separate retries, failed calls and streaming disconnects.
  3. Apply each provider’s exact input/output/cache rules and context tier.
  4. Apply Model.sale blended price to the measured charged-token total; the initial reservation is released after settlement.
  5. Compare the resulting total, latency, availability and model capability together.

Official reference rows below are informational, dated and linked to their source. They are not a promise of parity or a benchmark.

Official published rates

USD per million tokens, standard published rates. Cached input is shown separately where the source publishes it.

ModelInputCached inputOutputSource
OpenAI GPT-5.6 Sol$4.00See source$20.00Official pricing
OpenAI GPT-5.6 Terra$2.00See source$12.00Official pricing
OpenAI GPT-5.6 Luna$0.20See source$1.20Official pricing
Anthropic Claude Opus 4.8$5.00$0.50$25.00Official pricing
Anthropic Claude Sonnet 4.6$3.00$0.30$15.00Official pricing
Anthropic Claude Haiku 4.5$1.00$0.10$5.00Official pricing

Not an equivalence claim

Names, context limits and quality tiers can differ. Treat the table as a billing-shape reference.

Rates can change

Check the source date and verify official pages before committing to a workload.