Skip to content
InferenceHub

Documentation

InferenceHub is a byte-faithful gateway: one API key for Claude, GPT, and every model, served on their native APIs with 0% per-token markup. Point any compatible client at the gateway and keep your existing code.

Using an AI agent? Read these docs as raw markdown or point it at our llms.txt.

How it works

  1. Create a key in the portal — every account starts with $1 in free credit, no card required.
  2. Point your client at the gateway — it speaks the Anthropic and OpenAI wire formats your tools already use.
  3. Pick any model with the usual model field — one key reaches the whole catalog at native per-token rates.

Three minutes end to end — the quickstart has copy-paste requests for both wire formats.

Connect your coding agent

Each agent has a dedicated page with the exact config, which models it reaches, and its known gotchas:

Explore