Skip to main content

A/B testing models

Not sure whether tycho-1.0, clarke-1.0, or hal-1.0 is right for your use case? Run the same prompt through multiple models and score the outputs side by side.

The script

ab_test.py

Run it

What it does

  1. Runs the same prompt through tycho-1.0, clarke-1.0, and hal-1.0 concurrently.
  2. Measures latency precisely with time.perf_counter().
  3. Records token usage for cost comparison.
  4. Scores outputs with custom heuristics — length, formatting quality, or anything else you define.
  5. Prints a side-by-side comparison table with latency, tokens, and truncated responses.

Adding an LLM-as-judge scorer

Use a stronger model to evaluate outputs from weaker ones:

Tips

  • Run at least 10 trials per model — single-shot comparisons can be noisy.
  • Score on what matters to your product: accuracy, conciseness, latency, cost, or safety.
  • Keep a golden dataset of prompts you run on every model update to catch regressions.
  • Store results in a spreadsheet or DB so you can trend quality over time.

Next steps

Switching with fallback

Route requests to the winning model in production

Monitoring usage

Track quality and cost metrics over time