Which model here is cheapest, and what does each one cost?

11 models are published here with live token prices, across 2 routing prefixes, and every rate is the provider's own with no margin added. The cheapest input rate is Gemini 3.7 Flash (High) at $0.75 per million input tokens and $3.75 per million output; the most expensive is Claude Fable 5 at $10.00 input. Context windows are published per model in the table below. Prices are per million tokens in USD and are read from the upstream catalog twice an hour, so the table below is the authoritative figure rather than anything quoted elsewhere. These figures were last read from upstream on 2026-09-16. Credit is prepaid: what you top up becomes the spending ceiling on your API key, and one balance pays for any mixture of these models through a single OpenAI-compatible endpoint. There is no subscription, no seat count and no monthly minimum, so a dollar of credit buys a dollar of model usage whichever row you pick.

ModelModel idRouteContextInput / 1MOutput / 1M
Gemini 3.7 Flash (High) · 3 effort tiersantigravity/gemini-3.7-flash-highantigravityNot published$0.75$3.75
Claude 4.5 Haikucc/claude-haiku-4-5-20251001ccNot published$1.00$5.00
Claude Sonnet 5cc/claude-sonnet-5ccNot published$2.00$10.00
Claude 4.6 Sonnetcc/claude-sonnet-4-6ccNot published$3.00$15.00
Claude 4.5 Sonnetcc/claude-sonnet-4-5-20250929ccNot published$3.00$15.00
Claude Opus 5cc/claude-opus-5ccNot published$5.00$25.00
Claude Opus 4.8cc/claude-opus-4-8ccNot published$5.00$25.00
Claude Opus 4.7cc/claude-opus-4-7ccNot published$5.00$25.00
Claude Opus 4.6cc/claude-opus-4-6ccNot published$5.00$25.00
Claude Opus 4.5cc/claude-opus-4-5-20251101ccNot published$5.00$25.00
Claude Fable 5cc/claude-fable-5ccNot published$10.00$50.00
DevGPT

AI Model Pricing and Token Rates Catalog Guide

DevGPT passes model pricing straight through with zero markup added on top. Your credit is prepaid: the balance you top up becomes the exact spending ceiling on your API key. No subscriptions. No seat fees. No minimum spend requirements. Direct passthrough pricing.

We publish every rate openly without requiring an account.

We list exact input and output rates per million tokens across offered routes — including models from OpenAI, Anthropic, and Google under cc/, antigravity/, agy/, and cx/ prefixes. You can inspect context windows up to 1,048,576 tokens, maximum completion limits, and capability flags for vision, tool calling, reasoning, and thinking. Because we operate on a direct passthrough model, you pay only for the exact tokens your code consumes during inference.

Model IDs match upstream API specifications exactly. Point your OpenAI or Anthropic SDK code to our endpoint, drop in your key, and start running completions. The live table below refreshes twice every hour directly from upstream provider status.

Our cron reconciles token usage against your prepaid balance on regular ticks. You can monitor per-model cost breakdowns, request counts, and spend trends in your dashboard.

Inference calls connect directly to upstream provider endpoints, so this website never sits in your request path. Upstream enforces your spend limit, and we absorb any slight overshoot while an active request finishes so your balance never goes negative.

You can filter the catalog by connection type, search for specific model capabilities, or sort by token pricing directly in your browser. All published rates, context metrics, and provider availability flags refresh continuously.