> ## Documentation Index
> Fetch the complete documentation index at: https://docs.radium.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Introduction

> Production AI inference through OpenAI- and Anthropic-compatible endpoints

# Introduction

Radium serves the same class of models as the frontier labs on infrastructure it owns and operates. Point any OpenAI or Anthropic SDK at the Radium base URL, and three values change: base URL, API key and model string.

## Make your first request

<CodeGroup>
  ```bash cURL theme={null}
  curl -X POST https://api.radium.cloud/v1/chat/completions \\
    -H "Authorization: Bearer $RADIUM_API_KEY" \\
    -H "Content-Type: application/json" \\
    -d '{
      "model": "clarke-1.0",
      "messages": [{"role": "user", "content": "Write a rate limiter in TypeScript."}]
    }'
  ```

  ```python Python theme={null}
  from openai import OpenAI
  client = OpenAI(
      api_key="YOUR_RADIUM_API_KEY",
      base_url="https://api.radium.cloud/v1"
  )
  chat = client.chat.completions.create(
      model="clarke-1.0",
      messages=[{"role": "user", "content": "Write a rate limiter in TypeScript."}],
  )
  print(chat.choices[0].message.content)
  ```

  ```typescript TypeScript theme={null}
  import OpenAI from "openai";
  const client = new OpenAI({
    apiKey: process.env.RADIUM_API_KEY,
    baseURL: "https://api.radium.cloud/v1",
  });
  const chat = await client.chat.completions.create({
    model: "clarke-1.0",
    messages: [{ role: "user", content: "Write a rate limiter in TypeScript." }],
  });
  console.log(chat.choices[0].message.content);
  ```
</CodeGroup>

<Button href="https://deploy.radium.cloud/keys">Get an API key</Button>

## What changes when you switch

```diff TypeScript before-after theme={null}
  import OpenAI from 'openai';
  const client = new OpenAI({
-   apiKey: process.env.OPENAI_API_KEY,
-   baseURL: 'https://api.openai.com/v1',
+   apiKey: process.env.RADIUM_API_KEY,
+   baseURL: 'https://api.radium.cloud/v1',
  });
  const chat = await client.chat.completions.create({
-   model: 'gpt-4o',
+   model: 'clarke-1.0',
    messages: [{ role: 'user', content: 'Hello' }],
  });
```

## Models

| Model | Class | Best for | Context |
| - | - | - | - |
| `hal-1.0` | Opus-class | Complex reasoning, multi-step agents | 250,000 tokens |
| `clarke-1.0` | Sonnet-class | RAG, copilots, production default | 1,000,000 tokens |
| `tycho-1.0` | Haiku-class | Classification, extraction, routing | 125,000 tokens |

<Note>
  Pricing is dynamic. For current rates see [radium.cloud/pricing](https://radium.cloud/pricing).
</Note>

## OpenAI compatibility: how it works

Radium implements the standard OpenAI Chat Completions API. Every parameter you already use is supported, and the response shape is identical. Streaming, tool calling and JSON mode all work the same way. Your existing evals, prompt templates and middleware stay unchanged.

[Full guide](/switching/from-openai)

## Anthropic compatibility: how it works

Radium serves the Anthropic Messages endpoint at `POST /v1/messages`. Claude Code works with four environment variables. The Anthropic SDK, PydanticAI, and any framework that targets Anthropic can point at Radium by changing the base URL.

[Full guide](/switching/from-anthropic)

## Choosing a tier

Hal 1.0 is for reasoning: agents that decompose problems, generate structured plans, or write long programs. Use it when the task pays for quality. Clarke 1.0 is the default for most workloads: RAG, copilots, tool calls and any system where latency and cost matter. Tycho 1.0 is for high-volume automation: classification, entity extraction, summarization, routing. Send a request to Tycho when you would have used a classifier model or a rules engine.

[Models overview](/models/overview)

## First request

Export your key and send a request:

```bash theme={null}
export RADIUM_API_KEY="YOUR_RADIUM_API_KEY"
curl -X POST https://api.radium.cloud/v1/chat/completions -H "Authorization: Bearer $RADIUM_API_KEY" -H "Content-Type: application/json" -d '{"model": "clarke-1.0", "messages": [{"role": "user", "content": "Hello"}]}'
```

Or configure Claude Code:

```bash theme={null}
export ANTHROPIC_BASE_URL="https://api.radium.cloud"
export ANTHROPIC_AUTH_TOKEN="YOUR_RADIUM_API_KEY"
export ANTHROPIC_MODEL="hal-1.0"
export ANTHROPIC_SMALL_FAST_MODEL="tycho-1.0"
claude
```

## Common gotchas

**Wrong base URL**
OpenAI format requires `/v1` on the end of the base URL. Anthropic format does not. If requests return `404`, check which format you are using and whether `/v1` is present or absent.

**Wrong model string**
Using `gpt-*` or `claude-*` at Radium returns a model-not-found error. Use `hal-1.0`, `clarke-1.0` or `tycho-1.0`.

**Reading `content[0]` when reasoning is enabled**
When reasoning is on, the response contains a `thinking` content block before the `text` block. Read the last content block, not the first one. This is the same pattern the Anthropic SDK uses.

**Rate limits during load tests**
Load-testing against production endpoints will hit rate limits. Use the token-per-minute and request-per-minute values in the response headers to pace your tests, or contact sales for a higher limit.

## Next steps

<CardGroup cols={2}>
  <Card title="Quickstart" icon="bolt" href="/quickstart">Your first request, step by step</Card>
  <Card title="Models" icon="brain" href="/models/overview">Choose the right tier</Card>
  <Card title="From OpenAI" icon="arrows-rotate" href="/switching/from-openai">Migration guide and parameter map</Card>
  <Card title="From Anthropic" icon="arrows-rotate" href="/switching/from-anthropic">Messages API compatibility</Card>
  <Card title="Errors" icon="triangle-exclamation" href="/core-concepts/errors">Status codes and how to handle them</Card>
  <Card title="Rate limits" icon="gauge-high" href="/core-concepts/rate-limits">Limits and how to request more</Card>
</CardGroup>

## Enterprise

SSO/SAML, RBAC and audit logging. Uptime SLA available on request.

[Contact sales](https://radium.cloud/contact).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.