LLM evaluation

Detect a quality regression between models

Hard120 pts~45 min
  • Regression testing
  • Model comparison
Practice app · Acme Support Assistant

A deterministic LLM-style support assistant with retrieval (RAG), JSON mode, safety policies and tool calls, exposed via UI and API.

BASE_URL
/api/practice
Console app
/lab/ai-testing-detect-a-quality-regression-between-models

Your starter code already declares BASE_URL — call the API relative to it.

Objective

Run the golden set against both models and identify exactly which cases regress in acme-assistant-2.

Your task

  1. 1Run every eval case against acme-assistant-1 and acme-assistant-2 at temperature 0.
  2. 2Compute a pass rate per model and the list of cases that pass on -1 but fail on -2.
  3. 3Assert the pass rate of -2 is lower and the regressed ids are the shipping cases (standard-shipping, express-shipping).

Acceptance criteria

  • GET /ai/evals returns 200
  • POST /ai/chat returns 200
  • Both models are evaluated
  • At least 2 assertions pass

LLM evaluation · AI Testing · Evals & regression