RAG chatbot
Retrieval-Augmented Generation (RAG) grounds your chatbot in documents it can search at query time. This example uses an in-memory vector store, but the pattern works with any retrieval backend.Setup
pip install openai numpy
The full RAG pipeline
rag.py
import os
import numpy as np
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RADIUM_API_KEY"],
base_url="https://api.radium.cloud/v1",
)
# 1. Your knowledge base
documents = [
"Radium serves frontier-class models through OpenAI- and Anthropic-compatible endpoints.",
"Clarke is the default model for RAG, copilots, and production workloads.",
"Hal is the reasoning model for complex multi-step agents and code generation.",
"Tycho is the fast, cost-effective model for classification and extraction.",
"All Radium models support streaming, tool calling, and structured outputs.",
]
# 2. Simple in-memory vector store
class VectorStore:
def __init__(self):
self.documents = []
self.embeddings = []
def add(self, text: str):
response = client.embeddings.create(
model="clarke-1.0", # or your preferred embedding model
input=text,
)
embedding = np.array(response.data[0].embedding)
self.documents.append(text)
self.embeddings.append(embedding)
def search(self, query: str, top_k: int = 2) -> list[str]:
response = client.embeddings.create(
model="clarke-1.0",
input=query,
)
query_embedding = np.array(response.data[0].embedding)
# Cosine similarity
embeddings_matrix = np.stack(self.embeddings)
similarities = embeddings_matrix @ query_embedding
top_indices = np.argsort(similarities)[-top_k:][::-1]
return [self.documents[i] for i in top_indices]
# 3. Build the index
store = VectorStore()
for doc in documents:
store.add(doc)
print(f"Indexed {len(documents)} documents.\n")
# 4. Chat with retrieval
def chat(query: str) -> str:
# Retrieve relevant context
context = store.search(query, top_k=2)
context_block = "\n\n".join(f"- {d}" for d in context)
# Build the prompt
messages = [
{
"role": "system",
"content": (
"You are a helpful assistant. Use the provided context to answer. "
"If the context does not contain the answer, say so."
),
},
{
"role": "user",
"content": f"Context:\n{context_block}\n\nQuestion: {query}",
},
]
response = client.chat.completions.create(
model="clarke-1.0",
messages=messages,
temperature=0.3,
max_tokens=512,
)
return response.choices[0].message.content
# 5. Ask questions
if __name__ == "__main__":
questions = [
"Which model should I use for a coding assistant?",
"What features do all Radium models support?",
"Does Radium support image generation?",
]
for q in questions:
print(f"Q: {q}")
print(f"A: {chat(q)}\n")
Expected output
Indexed 5 documents.
Q: Which model should I use for a coding assistant?
A: Hal is the reasoning model for complex multi-step agents and code generation, making it the best choice for a coding assistant.
Q: What features do all Radium models support?
A: All Radium models support streaming, tool calling, and structured outputs.
Q: Does Radium support image generation?
A: The provided context does not mention image generation.
Production considerations
| Component | In-memory (demo) | Production |
|---|---|---|
| Vector store | numpy arrays | Pinecone, Weaviate, Qdrant, pgvector |
| Embeddings | Radium API | Dedicated embedding model or service |
| Document ingestion | Static list | Crawler + chunking pipeline |
| Re-ranking | None | Cross-encoder re-ranker |
| Caching | None | Redis for embedding cache |
Streaming RAG responses
Addstream=True to the chat completion and iterate over chunks to show the answer as it generates:
stream = client.chat.completions.create(
model="clarke-1.0",
messages=messages,
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="", flush=True)
print()
Next steps
Streaming response
Stream tokens in real time
Tool calling agent
Give your chatbot tools to act