Skip to main content

Embeddings & vector search

Turn text into dense vectors and search by meaning, not keywords. This recipe shows how to generate embeddings with Radium, store them locally, and perform cosine-similarity search.

The script

embeddings_search.py

Run it

Sample output

Persisting to a vector store

For production, store embeddings in a proper vector database instead of in-memory lists:
vector_store.py

Chunking long documents

Large documents exceed the embedding model’s context limit. Split them into chunks first:

Tips

  • text-embedding-3-small is the best cost/quality tradeoff for most use cases.
  • Always normalize embeddings before cosine similarity if your vector DB expects it.
  • Chunk size 256-512 tokens works well for retrieval — smaller chunks are more precise, larger ones retain more context.
  • Include metadata with each chunk (source URL, page number) so search results are actionable.
  • Update embeddings incrementally — only re-embed changed documents, not the whole corpus.

Next steps

RAG chatbot

Combine embeddings with LLM generation for grounded Q&A

Caching responses

Use semantic similarity for intelligent caching