Reliability methodology · public facts

Only measured availability is published.

Model.sale publishes a model as live only when a recent compatibility check, usage signal and active price agree. This page explains the checks so a developer or search system can interpret a status correctly.

Protocol checkA synthetic request validates JSON status, response shape, request ID and terminal usage. Streaming checks verify first bytes, event ordering and a terminal event.
PerformanceWe measure connect time, time to first byte, first token and total latency. Public pages show only rounded, non-customer aggregate facts when volume is sufficient.
Billing gateUsage must be present or covered by a documented deterministic policy. Price and margin are evaluated before publication; a health check alone cannot publish a model.
Current validation snapshot

6 live · 32 temporarily unavailable

Generated from the same public registry used by the API catalog. Last observed catalog check: 2026-09-12T09:30:59.660Z. This is a point-in-time compatibility snapshot, not a historical uptime percentage.

ModelStatusHealthSynthetic latencyChecked
claude-haiku-4-5UNAVAILABLEunhealthy · 0pending
claude-opus-4-5UNAVAILABLEunhealthy · 0pending
claude-opus-4-6UNAVAILABLEunhealthy · 0pending
claude-opus-4-7UNAVAILABLEunhealthy · 0pending
claude-opus-4-8UNAVAILABLEunhealthy · 0pending
claude-sonnet-4-5UNAVAILABLEunhealthy · 0pending
claude-sonnet-4-6UNAVAILABLEunhealthy · 0pending
deepseek-v4-flashLIVEhealthy · 1006838 ms2026-09-12T09:28:29.292Z
deepseek-v4-proLIVEhealthy · 1003688 ms2026-09-12T09:28:29.822Z
glm-5UNAVAILABLEunhealthy · 0pending
glm-5.1UNAVAILABLEunhealthy · 0pending
glm-5.2UNAVAILABLEunhealthy · 507553 ms2026-09-12T09:28:07.989Z
glm-5.3UNAVAILABLEunhealthy · 507026 ms2026-09-12T09:28:50.456Z
glm-5.3-flashUNAVAILABLEunhealthy · 0pending
gpt-5.4UNAVAILABLEunhealthy · 0pending
gpt-5.4-20UNAVAILABLEunhealthy · 0pending
gpt-5.4-miniUNAVAILABLEunhealthy · 50735 ms2026-09-12T09:28:56.697Z
gpt-5.5LIVEhealthy · 1002254 ms2026-09-12T09:28:52.362Z
gpt-5.6-lunaLIVEhealthy · 1002807 ms2026-09-12T09:29:19.393Z
gpt-5.6-solLIVEhealthy · 1007571 ms2026-09-12T09:29:27.685Z
gpt-5.6-terraLIVEhealthy · 10010180 ms2026-09-12T09:30:59.660Z
gpt-6-astraUNAVAILABLEunhealthy · 0pending
gpt-image-2UNAVAILABLEunhealthy · 0pending
kimi-k2.5UNAVAILABLEunhealthy · 0pending
kimi-k2.6UNAVAILABLEunhealthy · 0pending
kimi-k3UNAVAILABLEunhealthy · 0pending
mimo-v2-omniUNAVAILABLEunhealthy · 0pending
mimo-v2-proUNAVAILABLEunhealthy · 0pending
mimo-v2.5UNAVAILABLEunhealthy · 0pending
mimo-v2.5-proUNAVAILABLEunhealthy · 0pending
minimax-m2.5UNAVAILABLEunhealthy · 0pending
minimax-m2.7UNAVAILABLEunhealthy · 0pending
minimax-m3UNAVAILABLEunhealthy · 0pending
qwen-3.8UNAVAILABLEunhealthy · 0pending
qwen3.5-plusUNAVAILABLEunhealthy · 0pending
qwen3.6-plusUNAVAILABLEunhealthy · 0pending
qwen3.7-maxUNAVAILABLEunhealthy · 0pending
qwen3.7-plusUNAVAILABLEunhealthy · 0pending

What the health status means

StatusScoreMeaning
Healthy95–100Recent successful checks with normal latency and stream completion.
Degraded85–94Serving may continue while latency, 429s or synthetic results need attention.
UnhealthyBelow 85Customer publication is blocked or the model is temporarily unavailable.

The score is a compact operational signal, not a quality benchmark. It combines success, latency, streaming completion, rate-limit responses and synthetic checks.

Publication and refresh schedule

  1. Lightweight probes run every five minutes for protocol and endpoint health.
  2. A round-robin model sweep runs every eight hours; each result has a visible checked timestamp.
  3. A model stays public only while its check is within the freshness window, price is active and margin is at least 50%.
  4. After a failure, the catalog marks the exact model unavailable; it does not reroute the request to another model.
  5. Recovery requires another successful check and an audited publication action.

Privacy of measurement

Probes use synthetic prompts and never copy customer requests into shadow traffic. We retain status, latency, token counts and safe error classes for operations. Prompts, source code and response bodies are not written to logs, traces, analytics or Sentry.

For current incidents and component availability, see status.model.sale. For price calculation, see the pricing methodology.