Skip to main content

Introduction

Radium serves the same class of models as the frontier labs on infrastructure it owns and operates. Point any OpenAI or Anthropic SDK at the Radium base URL, and three values change: base URL, API key and model string.

Make your first request

What changes when you switch

TypeScript before-after

Models

Pricing is dynamic. For current rates see radium.cloud/pricing.

OpenAI compatibility: how it works

Radium implements the standard OpenAI Chat Completions API. Every parameter you already use is supported, and the response shape is identical. Streaming, tool calling and JSON mode all work the same way. Your existing evals, prompt templates and middleware stay unchanged. Full guide

Anthropic compatibility: how it works

Radium serves the Anthropic Messages endpoint at POST /v1/messages. Claude Code works with four environment variables. The Anthropic SDK, PydanticAI, and any framework that targets Anthropic can point at Radium by changing the base URL. Full guide

Choosing a tier

Hal 1.0 is for reasoning: agents that decompose problems, generate structured plans, or write long programs. Use it when the task pays for quality. Clarke 1.0 is the default for most workloads: RAG, copilots, tool calls and any system where latency and cost matter. Tycho 1.0 is for high-volume automation: classification, entity extraction, summarization, routing. Send a request to Tycho when you would have used a classifier model or a rules engine. Models overview

First request

Export your key and send a request:
Or configure Claude Code:

Common gotchas

Wrong base URL OpenAI format requires /v1 on the end of the base URL. Anthropic format does not. If requests return 404, check which format you are using and whether /v1 is present or absent. Wrong model string Using gpt-* or claude-* at Radium returns a model-not-found error. Use hal-1.0, clarke-1.0 or tycho-1.0. Reading content[0] when reasoning is enabled When reasoning is on, the response contains a thinking content block before the text block. Read the last content block, not the first one. This is the same pattern the Anthropic SDK uses. Rate limits during load tests Load-testing against production endpoints will hit rate limits. Use the token-per-minute and request-per-minute values in the response headers to pace your tests, or contact sales for a higher limit.

Next steps

Quickstart

Your first request, step by step

Models

Choose the right tier

From OpenAI

Migration guide and parameter map

From Anthropic

Messages API compatibility

Errors

Status codes and how to handle them

Rate limits

Limits and how to request more

Enterprise

SSO/SAML, RBAC and audit logging. Uptime SLA available on request. Contact sales.