Context
131.1K
Input
$1.50
Output
$7.50
Cache read
$0.15
Cache write
-
Context
131.1K
Input
$1.50
Output
$7.50
Cache read
$0.15
Cache write
-
Context
524.3K
Input
$0.68
Output
$2.09
Cache read
$0.07
Cache write
-
Context
1.0M
Input
$0.3
Output
$1.20
Cache read
$0.006
Cache write
-
Context
1.1M
Input
$2
Output
$10
Cache read
$0.1
Cache write
$2.50
Context
1.1M
Input
$4
Output
$20
Cache read
$0.2
Cache write
$5
Context
1M
Input
$2
Output
$10
Cache read
$0.2
Cache write
$2.50
Context
1.0M
Input
$0.3
Output
$1.20
Cache read
$0.006
Cache write
-
Context
1M
Input
$4
Output
$12
Cache read
$0.5
Cache write
$5
Context
1.0M
Input
$3
Output
$15
Cache read
$0.3
Cache write
-
Context
1M
Input
$4
Output
$20
Cache read
$0.2
Cache write
$5
Context
1M
Input
$8
Output
$40
Cache read
$0.4
Cache write
$10
Context
1.1M
Input
$0.1
Output
$0.5
Cache read
$0.01
Cache write
$0.125
Context
1.1M
Input
$0.2
Output
$1
Cache read
$0.02
Cache write
$0.25
Context
1.1M
Input
$2
Output
$10
Cache read
$0.2
Cache write
$2.50
Context
1.1M
Input
$4
Output
$20
Cache read
$0.4
Cache write
$5
Context
500K
Input
$2
Output
$6
Cache read
$0.5
Cache write
-
Context
1.0M
Input
$0.04
Output
$1.28
Cache read
$0.04
Cache write
-
Context
1.0M
Input
$0.435
Output
$0.87
Cache read
$0.004
Cache write
-
Context
1.0M
Input
$4.35
Output
$8.70
Cache read
$0.036
Cache write
-
Context
1M
Input
$1
Output
$2.70
Cache read
$0.05
Cache write
-
Context
1M
Input
$0.37
Output
$1.25
Cache read
$0.075
Cache write
-
Context
1M
Input
$0.15
Output
$0.47
Cache read
$0.016
Cache write
-
Context
131.1K
Input
$4
Output
$20
Cache read
$0.4
Cache write
$5
Context
131.1K
Input
$6
Output
$30
Cache read
$0.6
Cache write
$7.50
Context
1M
Input
$2
Output
$6
Cache read
$0.25
Cache write
-
1–25 of 257
Token prices are per 1M tokens; expand any model for full rates and capabilities, or open its detail page for the quickstart. The filters, sorting, search, pagination, and compare flow match the primary API catalog. Models marked DEDICATED are open-weight and available through Enterprise Inference as a single-tenant deployment.
Most models in this catalog run on the infrastructure of other providers; the Router is how you reach them safely. The models marked DEDICATED are open-weight, and Enterprise Inference runs them single-tenant on capacity reserved for you.
Every model in this catalog through one OpenAI-compatible endpoint, billed per token at the rates listed above. The Router reaches only upstream providers that are Zero Data Retention and do not train on your data.
Models marked DEDICATED are open-weight and can run in a deployment that is only yours: single-tenant, on capacity reserved for you, isolated from every other customer.
Every model above is one request away. OpenAI-compatible, with intelligent routing and prompt caching included.