Three different products
Together AI documents serverless per-token inference for open-weight models and separate dedicated or provisioned capacity for predictable throughput. Fireworks combines serverless per-token inference with GPU-time pricing for dedicated deployments. Groq focuses on a curated, low-latency catalog with production and preview model tiers.
A prepaid gateway such as Model.sale adds a different layer: one wallet, a curated public catalog and request-level settlement. It can be convenient for a small team, but it should not be confused with owning capacity at the underlying inference platform.
Together AI: broad open-model choice
Together's public pricing page lists model-specific input, cached-input and output rates. Its serverless documentation describes pay-per-token usage with no provisioning cost or minimum, while dedicated products trade flexibility for reserved capacity.
This is a strong fit when you want direct access to open models and are comfortable choosing the model ID, reading its capability table and managing the provider account yourself.
Fireworks: serverless first, dedicated later
Fireworks describes serverless inference as per-token and offers dedicated deployments priced by GPU time. The two modes solve different problems: serverless is convenient for variable traffic, while dedicated capacity can make sense once utilization and latency are predictable.
When estimating cost, include idle or reserved GPU time for dedicated deployments. A low per-token serverless number cannot be compared directly with a GPU-hour quote.
Groq: optimize for latency and a focused catalog
Groq publishes production and preview model tables with token rates and model-specific limits. Its developer billing is postpaid with progressive thresholds, which is materially different from a prepaid wallet.
Groq can be compelling for latency-sensitive workloads that fit its active catalog. It is not a general multi-provider aggregator, so you should plan what happens when a required model is not listed or a preview model is retired.
The five numbers to record
For every candidate, record (1) input and output price, (2) cache and reasoning rules, (3) payment timing and fees, (4) rate limits and concurrency, and (5) P95 time-to-first-token plus stream failure rate from your own synthetic checks.
Then run the same prompt set with the same token budget. Keep prompts synthetic and non-sensitive, store only redacted reports and repeat the test after a day; a single fast request is not a capacity benchmark.
A practical choice for coding agents
Use a direct open-model platform when you need provider-native features, model breadth or dedicated capacity. Use a prepaid gateway when the main job is giving a small team one budgeted key and a short list of verified models. Use a multi-provider router when fallback and model discovery justify the extra billing and policy complexity.
In all cases, publish the check date and source URL. Avoid blanket claims such as “90% cheaper” unless the workload, token mix, currency and fee treatment are shown next to the number.