# InferenceHub Blog

> Essays from the founders — the honest economics of frontier inference, how we price, and why.
> Newest first. Also rendered at https://inferencehub.tech/blog with RSS at https://inferencehub.tech/blog/rss.xml.

## We're not cheaper than the labs

*August 5, 2026 — Kevin, co-founder* · https://inferencehub.tech/blog/honest-pricing

A few times a week, someone lands in our Discord and asks a completely reasonable question: *another service sells "the same" Claude access at a third of your price — why would I pay list through you?*

This post is the answer we give them, written down. It costs us signups. We think it's also why the customers we do have will still be customers next quarter.

### The part nobody says out loud

Our wholesale cost for frontier models is roughly what the labs charge at list. Not a little below. Roughly at.

That isn't a failure of negotiation. Frontier inference is priced tightly because the labs can sell every token they can serve — real supply, bought through real enterprise channels, tracks list price. So when you run Claude or GPT through us, the honest price is parity: the lab's published rate, with 0% per-token markup. Our business runs on flat-fee subscriptions, not on a hidden token spread.

Which raises the obvious question: how is anyone selling real frontier tokens far below list, *permanently*? In our experience there are exactly three ways, and each one is a clock.

### Clock one: burning investor money

The oldest clock. A subsidized price is a marketing expense — somebody's funding round is paying the difference between what you're charged and what the tokens cost. It's a real discount while it lasts, and it never lasts. The subsidy ends the day the funding mood changes, and the price snaps back to list with a blog post about "aligning pricing with sustainable growth." If you've been in this industry more than a year, you've read that blog post several times.

How to spot it: the price is below list, the supply is legitimate, and the company's business model doesn't explain the gap.

### Clock two: reselling pooled logins

The gray-market clock. Consumer chat subscriptions — bought, borrowed, or shared — relayed through automation so that many customers ride accounts priced for one human. It's a violation of the terms those subscriptions are sold under, which is why this category lives in a permanent ban wave: accounts die, capacity vanishes, and your coding agent dies mid-session with them. The operations themselves rotate names after every wave, which tells you what their own life expectancy looks like from the inside.

How to spot it: frontier-model access at a fraction of list, paid to a brand that didn't exist two quarters ago, with capacity that comes and goes.

### Clock three: relabeling a cheaper model

The quiet clock. An open-weight model, remapped to wear a frontier name. Sometimes the substitute is genuinely good — open-weight models have gotten remarkably strong, and we sell plenty of them, honestly labeled, at honestly low prices. But a good model wearing the wrong name is still not the model on the label, and you're paying frontier prices for it. If you suspect this, diff the outputs on tasks you know well. The fingerprints show.

How to spot it: "Claude" that's oddly fast, oddly cheap, and oddly bad at the things Claude is good at.

### All three clocks run out

We want to be precise about the claim. We are not saying cheap access doesn't work — it often works fine, for months. We're saying it's running on a clock, and you don't get to see the dial. If price alone decides it for you, the relays are roughly 3× cheaper: go in clear-eyed about which clock you're on, and don't build a workflow you can't afford to have die on a Tuesday.

We'd rather be the boring option that's still here next quarter.

### Where that puts us

Picture the market on two axes: how fast you can start, and where the supply comes from. Cloud marketplaces and enterprise channels have real supply — behind the accounts, contracts, and procurement that come with them. Relays and shared subscriptions start instantly — on supply that can vanish mid-session. Nobody chooses onboarding friction *and* supply risk, so that corner is empty.

We hold the remaining corner: paid-for enterprise supply on each model's native wire, and you can sign up, top up by card or crypto — private by default — and point your agent at it in minutes. Cheaper than us usually means gray-market or relabeled. More enterprise than us usually means procurement. We're the corner with neither compromise.

### So what do you actually pay?

Pay-as-you-go is parity: the lab's list price, 0% per-token markup, no fee on top-ups. The business is the subscription ladder — All Access at $10/month, Plus at $25, Max at $50 — where every plan gets every model and every coding agent on one key: a pooled premium budget for flagship Claude and GPT, generous standard windows for everything else (the open-weight daily drivers and lower-priced frontier), and when you hit a window you keep working at pay-as-you-go rates. No hard stop, no wall at 2 a.m.

Two things we'll volunteer before you ask. First, in-window subscription usage routinely costs us more than the flat fee — we fund that gap deliberately, as our customer-acquisition spend, and it's bounded by the windows we publish rather than by fine print we don't. Second, we're not the right choice for everyone: if you need procurement-grade contracts and SLAs, the enterprise channels are the honest recommendation; if a relay's risk profile is fine for your use, they are genuinely cheaper.

Everything else about us assumes you'd rather be told this than discover it. There's $1 of free credit on signup — enough to point Claude Code, Codex, or Cursor at us and watch real Opus and Sonnet calls stream back before you pay anything. And your first month on any plan is matched 100% in credit right now, which is the closest thing to a discount we'll ever run: it's dev credit on top, never a lower sticker.

*— Kevin, co-founder, InferenceHub*
