Introduction
Radium serves the same class of models as the frontier labs on infrastructure it owns and operates. Point any OpenAI or Anthropic SDK at the Radium base URL, and three values change: base URL, API key and model string.Make your first request
What changes when you switch
TypeScript before-after
Models
Pricing is dynamic. For current rates see radium.cloud/pricing.
OpenAI compatibility: how it works
Radium implements the standard OpenAI Chat Completions API. Every parameter you already use is supported, and the response shape is identical. Streaming, tool calling and JSON mode all work the same way. Your existing evals, prompt templates and middleware stay unchanged. Full guideAnthropic compatibility: how it works
Radium serves the Anthropic Messages endpoint atPOST /v1/messages. Claude Code works with four environment variables. The Anthropic SDK, PydanticAI, and any framework that targets Anthropic can point at Radium by changing the base URL.
Full guide
Choosing a tier
Hal 1.0 is for reasoning: agents that decompose problems, generate structured plans, or write long programs. Use it when the task pays for quality. Clarke 1.0 is the default for most workloads: RAG, copilots, tool calls and any system where latency and cost matter. Tycho 1.0 is for high-volume automation: classification, entity extraction, summarization, routing. Send a request to Tycho when you would have used a classifier model or a rules engine. Models overviewFirst request
Export your key and send a request:Common gotchas
Wrong base URL OpenAI format requires/v1 on the end of the base URL. Anthropic format does not. If requests return 404, check which format you are using and whether /v1 is present or absent.
Wrong model string
Using gpt-* or claude-* at Radium returns a model-not-found error. Use hal-1.0, clarke-1.0 or tycho-1.0.
Reading content[0] when reasoning is enabled
When reasoning is on, the response contains a thinking content block before the text block. Read the last content block, not the first one. This is the same pattern the Anthropic SDK uses.
Rate limits during load tests
Load-testing against production endpoints will hit rate limits. Use the token-per-minute and request-per-minute values in the response headers to pace your tests, or contact sales for a higher limit.
Next steps
Quickstart
Your first request, step by step
Models
Choose the right tier
From OpenAI
Migration guide and parameter map
From Anthropic
Messages API compatibility
Errors
Status codes and how to handle them
Rate limits
Limits and how to request more