uptime 99.994%/28 models · 10 providers

One key.
Every model.
Zero rate ceilings.

Forge is an OpenAI-compatible routing gateway. Point your SDK at us and stream from GPT-5.6, Claude Opus, Grok, DeepSeek, Kimi and 20+ more — with automatic key pooling and millisecond failover.

22.1B
tokens routed
573ms
avg first-token
99.99%
routing SLA
gateway.forge/edge · iad1
live
21:02:18200POST /v1/chat/completionsclaude-sonnet-4-5iad1214ms
21:02:16200POST /v1/chat/completionsdeepseek-r1sfo1891ms
21:02:15200POST /v1/chat/completionsgpt-5.6-lunafra1342ms
21:02:14200GET /v1/models-iad112ms
21:02:12200POST /v1/chat/completionsgrok-4.5iad1512ms
21:02:11429POST /v1/chat/completionskimi-k3hkg11024ms
21:02:10200POST /v1/chat/completionskimi-k3hkg163ms
21:02:09200POST /v1/chat/completionsgemini-3-proiad1421ms
21:02:07200POST /v1/embeddingstext-embed-3iad187ms
21:02:06200POST /v1/chat/completionsclaude-sonnet-4-5iad1214ms
21:02:05200POST /v1/chat/completionsdeepseek-r1sfo1891ms
21:02:03200POST /v1/chat/completionsgpt-5.6-lunafra1342ms
21:02:02200GET /v1/models-iad112ms
21:02:01200POST /v1/chat/completionsgrok-4.5iad1512ms
21:01:59429POST /v1/chat/completionskimi-k3hkg11024ms
21:01:58200POST /v1/chat/completionskimi-k3hkg163ms
21:01:57200POST /v1/chat/completionsgemini-3-proiad1421ms
21:01:56200POST /v1/embeddingstext-embed-3iad187ms
Routing toanthropicopenaigooglexaimetamoonshotglmdeepseekminimaxqwentencentxiaomizinccustom
/* 01 — integration */

Drop-in for the OpenAI SDK.

Change two lines. Keep your streaming, tool-calls, JSON mode. Every request is observable in the Forge dashboard and billed against a single balance.

base_url
https://run.forgeapi.org/v1
provider proxy
OrcaRouter Gateway (api.orcarouter.ai)
auth
Bearer fg-*
endpoints
/chat/completions · /embeddings · /models
region
iad1 · sfo1 · fra1 · hkg1
agent.pycurl.shnode.ts
from openai import OpenAI
client = OpenAI(
    api_key="fg-demotoken1234567890",
    base_url="https://run.forgeapi.org/v1",
)
response = client.chat.completions.create(
    model="deepseek-r1",
    messages=[{"role": "user", "content": "Explain load balancing."}],
    stream=True,
)
for chunk in response:
    print(chunk.choices[0].delta.content, end="")
streaming · 1.2s to first token · routed deepseek-r1 via iad1
/* 02 — catalog */

Frontier models, one endpoint.

full catalog →
AnthropicPRO
anthropic/claude-fable-5
1M ctx$12.50 / $62.50
AnthropicPRO
anthropic/claude-haiku-4.5
200K ctx$1.25 / $6.25
AnthropicPRO
anthropic/claude-opus-4.5
200K ctx$6.25 / $31.25
AnthropicPRO
anthropic/claude-opus-4.6
1M ctx$6.25 / $31.25
AnthropicPRO
anthropic/claude-opus-4.7
1M ctx$6.25 / $31.25
AnthropicPRO
anthropic/claude-opus-4.8
1M ctx$6.25 / $31.25
AnthropicPRO
anthropic/claude-opus-5
1M ctx$6.25 / $31.25
AnthropicPRO
anthropic/claude-sonnet-4.5
1M ctx$3.75 / $18.75
/* 03 — plans */

Pay for tokens, not seats.

No forced monthly seat minimums. Transparent usage billing with zero markup on upstream provider tokens.

Developer Pay-As-You-Go

POPULAR

Ideal for individual developers building agents, scripts, and production apps with transparent token billing.

$0per month
  • Access to all 28+ frontier models (Claude 3.7, GPT-4o, DeepSeek R1)
  • Zero monthly seat fees — pay only for tokens you consume
  • OpenAI SDK drop-in replacement (base_url: https://run.forgeapi.org/v1)
  • Automatic key rotation & sub-500ms edge failover
  • Real-time telemetry & latency metrics dashboard

Pro Scale

RECOMMENDED

Built for teams and heavy agent workloads requiring guaranteed high rate limits and dedicated throughput.

$10per month
  • $10.00 included monthly usage credit
  • Higher rate limit ceilings across all providers (100k+ TPM)
  • Bring-Your-Own-Key (BYOK) fallback integration
  • Team workspace & shared API key management
  • 99.99% Routing Uptime SLA guarantee
  • Priority Discord & Email developer support

Enterprise Custom

ENTERPRISE

Dedicated routing clusters, zero data retention, and custom SLAs for enterprise AI applications.

Customflexible billing
  • Dedicated isolated edge proxy clusters
  • Zero Data Retention (ZDR) & SOC-2 compliance
  • Custom SLA guarantees & custom rate limits
  • Custom model fine-tuning & private endpoint proxying
  • 24/7 Phone & Slack channel support
  • Custom invoice & purchase order billing