Skip to content
InferenceHub
Pricing

Same prices you know. Everything un-siloed.

InferenceHub charges the model provider’s native per-token rate with 0% markup. Pay as you go, or add a plan from $10/mo — either way one key runs every model and every coding agent, and you never hit a hard wall.

Launch offer: your first month matched 100% in credit.

How it compares

We don’t undercut the providers — we charge their rates. What you get for the same price is reach: every model, every agent, one key, no lock-in.

CapabilityInferenceHubClaude CodeCodexCursor
Models you can useClaude, GPT, GLM, DeepSeek, Kimi, Qwen & moreAnthropic models onlyOpenAI models onlyA curated set, inside the editor
Coding agents on one keyClaude Code, Codex, Cursor, Kilo, HermesClaude Code onlyCodex onlyCursor editor only
Switch model & tool without switching accounts
When you hit your limitKeep going at pay-as-you-go — no hard stopBlocked until the window resetsBlocked until the window resetsThrottled / blocked until reset
Deep Research includedIncluded — powered by GLM-5.2, doesn't touch your premium budgetCounts against your plan limitsCounts against your plan limits
Pay-as-you-go at native rates (no subscription)
Top up with crypto · no KYC

Comparison reflects each product’s standard plans as of July 2026. Per-token rates always match the native provider — no markup. Claude Code reaches Claude AND the open-weight daily drivers (GLM-5.2, DeepSeek, Kimi, MiniMax, Qwen); Cursor and other OpenAI-compatible editors reach the open-weight chat models AND Claude (via the gateway’s translation lane); Kilo Code reaches the full catalog — Claude, open-weight, and GPT-5.6.

Subscription

Every model. Three sizes.

Every plan gets every model and every coding agent on one key — a premium budget for flagship Claude and GPT, standard windows for everything else. Bigger plans buy bigger windows and bigger budgets.

All Access

every model, one key

$10/mo

Covers up to $300 of usage

30× the price, at native rates

Every model, including Opus — on one key.

  • $50/mo premium budget — Claude + GPT pooled, burstable, incl. Opus & GPT-5.6 Sol
  • Standard windows — $10 per 5-hour window, $60/week — for GLM-5.2, DeepSeek, Kimi, MiniMax, Qwen + lower-priced frontier like Sonnet
  • Deep Research included in hosted chat — powered by GLM-5.2
  • Every coding agent (Claude Code, Codex, Cursor, Kilo, Hermes, ZCode) — every open-weight model runs inside Claude Code
  • Overflow at 0%-markup pay-as-you-go — no hard stop

Premium models: One $50/month premium budget shared across Claude and GPT (Opus, GPT-5.6 Sol) — burstable, not weekly-paced, and pooled: spend all of it on either vendor. Every other model — open-weight and lower-priced frontier alike — draws from the much larger standard windows.

Most popular

Plus

for heavier daily driving

$25/mo

Covers up to $1,000 of usage

40× the price, at native rates

Bigger windows, serious frontier budgets.

  • $200/mo premium budget — Claude + GPT pooled, burstable, incl. Opus & GPT-5.6 Sol
  • Standard windows for multi-agent sessions — $40 per 5-hour window, $250/week
  • Everything in All Access — every model, every agent, Deep Research, no hard stop

Premium models: One $200/month premium budget shared across Claude and GPT (Opus, GPT-5.6 Sol) — burstable, not weekly-paced, and pooled: spend all of it on either vendor. Every other model — open-weight and lower-priced frontier alike — draws from the much larger standard windows.

Max

for all-day agent fleets

$50/mo

Covers up to $2,500 of usage

50× the price, at native rates

The heaviest windows we sell.

  • $500/mo premium budget — Claude + GPT pooled, burstable, incl. Opus & GPT-5.6 Sol
  • Standard windows for fleets — $100 per 5-hour window, $625/week
  • Everything in All Access — every model, every agent, Deep Research, no hard stop

Premium models: One $500/month premium budget shared across Claude and GPT (Opus, GPT-5.6 Sol) — burstable, not weekly-paced, and pooled: spend all of it on either vendor. The standard windows are sized for fleets of agents.

The premium budget covers premium models (Opus, GPT-5.6 Sol), pooled across Claude and GPT. It's burstable — spend it all on either vendor, in one day if you want. Every other model, open-weight and lower-priced frontier alike, draws the standard windows, which reset every 5 hours and weekly. Reach any limit and you keep working at pay-as-you-go rates — no hard stop. Pay by card, or by crypto (pay-per-period, no KYC).

Your first month's payment is matched 100% in bonus credit, usable for inference only and expiring 30 days after issue. New subscribers only, one per customer.

No commitment

Prefer to just pay as you go?

Skip the subscription entirely. Pay only for the tokens you use, at the model provider's native rate — the same price as going direct.

  • Native per-token rates — 0% markup, on every model
  • Start with $1 in free credit — no credit card required
  • No monthly commitment and no minimums
  • Top up by card or crypto — no lock-in, no KYC
  • Real-time usage and balance tracking in the portal
  • Subscriptions overflow here automatically — so you never hit a wall

Pricing questions

More in the FAQ and docs.

Yes. You pay the model provider's native per-token rate with 0% markup — the price in our Models list is the price you pay. We don't surcharge your tokens; subscriptions and scale are how the business works, not a hidden markup.

One key. Every model. No markup.

Get started in minutes with $1 in free credit — no credit card required.