Skip to content
Launch offer: your first month matched 100% in credit.

Every model. Every coding agent.
One key.

Claude, GPT, and open-weight models at native rates — 0% per-token markup.Plans from $10/mo, or just pay as you go. No hard stop.
// ~/.claude/settings.json
{
"env": {
"ANTHROPIC_BASE_URL": "https://app.inferencehub.tech",
"ANTHROPIC_AUTH_TOKEN": "sk-prov-live-YOUR_KEY",
"CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS": "1",
"CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY": "1"
},
"model": "sonnet"
}
ConnectedNative API0% markup
20+ models

Access the best models

Choose from the latest frontier and open-weight models.

More

Claude Fable 5

Anthropic

$10.00
/ 1M input tokens
$50.00
/ 1M output tokens
$1.00
/ 1M cached input tokens

GPT-5.6 Sol

OpenAI

$5.00
/ 1M input tokens
$30.00
/ 1M output tokens
$0.50
/ 1M cached input tokens

GLM 5.2

Z.ai

$1.05
/ 1M input tokens
$4.40
/ 1M output tokens
$0.21
/ 1M cached input tokens

DeepSeek V4 Pro

DeepSeek

$1.392
/ 1M input tokens
$2.784
/ 1M output tokens
$0.348
/ 1M cached input tokens

Kimi K3

Moonshot AI

$3.00
/ 1M input tokens
$15.00
/ 1M output tokens
$0.60
/ 1M cached input tokens
Built for developers

Works with your favorite AI coding agents

Use InferenceHub as the inference backend for all your tools and agents.

Desktop app

Every model, out of the browser.

A native chat app for macOS and Windows — web search, deep research, and project files, on the same one key.

Subscription

Every model. Three sizes.

Every plan gets every model and every coding agent on one key — a premium budget for flagship Claude and GPT, standard windows for everything else. Bigger plans buy bigger windows and bigger budgets.

Launch offer: your first month matched 100% in credit.

All Access

every model, one key

$10/mo

Every model, including Opus — on one key.

  • $50/mo premium budget — Claude + GPT pooled, burstable, incl. Opus & GPT-5.6 Sol
  • Standard windows — $10 per 5-hour window, $60/week — for GLM-5.2, DeepSeek, Kimi, MiniMax, Qwen + lower-priced frontier like Sonnet
  • Deep Research included in hosted chat — powered by GLM-5.2
  • Every coding agent (Claude Code, Codex, Cursor, Kilo, Hermes, ZCode) — every open-weight model runs inside Claude Code
  • Overflow at 0%-markup pay-as-you-go — no hard stop
Most popular

Plus

for heavier daily driving

$25/mo

Bigger windows, serious frontier budgets.

  • $200/mo premium budget — Claude + GPT pooled, burstable, incl. Opus & GPT-5.6 Sol
  • Standard windows for multi-agent sessions — $40 per 5-hour window, $250/week
  • Everything in All Access — every model, every agent, Deep Research, no hard stop

Max

for all-day agent fleets

$50/mo

The heaviest windows we sell.

  • $500/mo premium budget — Claude + GPT pooled, burstable, incl. Opus & GPT-5.6 Sol
  • Standard windows for fleets — $100 per 5-hour window, $625/week
  • Everything in All Access — every model, every agent, Deep Research, no hard stop

The premium budget covers premium models (Opus, GPT-5.6 Sol), pooled across Claude and GPT. It's burstable — spend it all on either vendor, in one day if you want. Every other model, open-weight and lower-priced frontier alike, draws the standard windows, which reset every 5 hours and weekly. Reach any limit and you keep working at pay-as-you-go rates — no hard stop. Pay by card, or by crypto (pay-per-period, no KYC).

Compare plans & see how we stack up

Your first month's payment is matched 100% in bonus credit, usable for inference only and expiring 30 days after issue. New subscribers only, one per customer.

No commitment

Or just pay as you go.

  • Same rates as native providers. No markup.
  • 100% usage-based — pay only for what you use
  • No monthly commitment; start with $1 free credit
  • Daily reward for active accounts — claimable credit that grows with your streak
  • Real-time usage and spend tracking
  • Top up by card or crypto — no lock-in
Limited time offer

Double your credits on your first top-up

Top up during the launch period and we’ll match it, 100%.

Launch offer: 100% match on your first top-up, up to $100 in bonus credit. Bonus credit is promotional — usable for inference only and expires after 30 days. Per-token rates always match the native provider — no markup.

Honest pricing

We're not cheaper than the labs.

Here's why developers route through us anyway.

Our wholesale cost for frontier models is roughly what the labs charge at list. So that's the price you pay — native rates, 0% per-token markup — and the business runs on subscriptions, not a hidden token spread. Anyone selling real frontier tokens far below list, permanently, is running one of three clocks:

Clock 1

Burning investor money

A subsidized price that snaps back to list the day the funding mood changes. Great while it lasts. It never lasts.

Clock 2

Reselling pooled logins

Shared or borrowed credentials relayed until the ban wave hits — and your coding agent dies mid-session with them.

Clock 3

Relabeling a cheaper model

An open-weight model remapped to wear a frontier name. Sometimes genuinely good — but not the model on the label. Diff the outputs.

All three clocks run out. If price alone decides it, the relays are ~3× cheaper — go in clear-eyed about why. We'd rather be the boring option that's still here next quarter.

Where that puts us

Legitimate supply · Start in minutes

InferenceHub

Paid-for enterprise supply on each model's native wire. Sign up, top up by card or crypto — private by default — and point your agent at it.

Legitimate supply · Procurement-grade onboarding

Cloud marketplaces

Real supply through hyperscaler and enterprise channels — with the accounts, contracts, and setup that come with them.

Gray-market supply · Start in minutes

Relays & shared subscriptions

Instant and ~3× cheaper — running on pooled credentials that can vanish mid-session, rotating names after every ban.

Gray-market supply · Procurement-grade onboarding

The empty corner

Nobody picks onboarding friction and supply risk at the same time. No one lives here.

Cheaper than us usually means gray-market or relabeled. More enterprise than us means procurement. We're the corner with neither compromise.

Daily reward

Show up daily, get credited daily

A small thank-you for active accounts — fixed amounts on a published ladder, no wheels to spin. Claim it in the portal's Rewards tab.

  • Use it, claim it

    Make at least one successful API request in a day (UTC) and a claim unlocks in your portal — $0.10 of open-weight credit, good for 48 hours.

  • Streaks grow it

    Claim on consecutive days and the daily credit ramps up to $0.25 by day seven. Miss a day and the streak simply restarts — nothing else is lost.

  • Milestones go frontier

    Every 7-day streak adds a $0.50 bonus you can spend on any model — Claude and GPT included. A 30-day streak bumps that to $2.00.

Why InferenceHub

Faithful, private, drop-in

  • Byte-faithful by design

    A pure passthrough to the upstream model — same tokens, same outputs, byte for byte. No silent re-tuning, no quality shortcuts.

  • Private, no account

    Top up with crypto, get a key, start calling. No signup, no KYC. We never train on or retain your prompts.

  • Honest parity pricing

    Native rates with no markup games or hidden routing fees. Crypto top-ups shave a little more off the top.

  • Drop-in compatible

    Native Anthropic & OpenAI APIs. Point your existing SDK or coding agent — Claude Code, Cursor, Cline — at our URL. Nothing else changes.

Ship AI features faster.

Get an API key in minutes and point your existing client at it. Start with $1 in credit — no credit card required.