Skip to content

Model catalogue & prices

Every model. One API.

Live prices per million tokens from the approved catalogue snapshot, capability evidence for every route, and a dated comparison against each developer's standard direct API list price. No hidden multipliers, no surprise postpaid bills.

Catalogue temporarily unavailable

Live prices cannot be shown right now.

InferencePass does not display stale or unverified rates as current. The approved catalogue source did not respond. Try again shortly or check the service status.

Next in the catalogue

Coming soon: four routes under evaluation.

These model identities and routes are being evaluated now. They are visible as a roadmap, but are not available for API traffic until verification, commercial approval, and a published price snapshot are complete.

  • OpenAI-compatible claimUnder evaluation

    GPT-6 Astra

    gpt-6-astra

    Frontier reasoning

    A frontier reasoning route being evaluated for long-horizon work and high output headroom.

  • Anthropic-compatible claimUnder evaluation

    Claude Opus 4.8

    claude-opus-4-8

    Premium reasoning

    A premium reasoning route under identity, capability, and commercial review.

  • DeepSeek-compatible claimUnder evaluation

    DeepSeek V4 Flash 0731

    deepseek-v4-flash-0731

    Fast reasoning value

    A low-cost fast reasoning route being checked for price and usage reconciliation.

  • Alibaba Qwen-compatible claimUnder evaluation

    Qwen 3.8 Flash

    qwen3.8-flash

    Fast multimodal value

    A fast multimodal route under regional pricing and capability review.

Questions

How the catalogue works.

How is InferencePass cheaper than direct API prices?

InferencePass is an independent gateway that routes requests over third-party Pollinations community routes and publishes per-model prices from its own approved price snapshot. Where unit economics allow, selected routes are offered below the dated standard direct API list prices of the named model developer, and every comparison on this page shows its reference date and source.

Are these the official APIs of OpenAI, Anthropic, or Google?

No. InferencePass is an independent service. Models are delivered through third-party routes that claim compatibility with the named models. InferencePass does not state or imply that any route is operated by, endorsed by, or cryptographically verified as the named developer's official API.

How am I billed for model usage?

You buy prepaid credits, create an OpenAI-compatible API key, and every request meters uncached input, cached input, cache writes, output, and reasoning tokens where the route reports them. Charges are deducted from your credit balance against a price snapshot accepted at request time, so prices never change retroactively for accepted requests.

What happens if a route has an outage?

Route health is monitored continuously. Requests use pre-approved fallback routes with the same verified model identity where one exists; otherwise they fail with a documented retryable error and undelivered usage is not billed. Current availability is shown on every model card and on the status page.