LLM evaluation

Score outputs with an LLM judge

Hard120 pts~45 min
  • LLM-as-judge
  • Rubrics
Practice app · Acme Support Assistant

A deterministic LLM-style support assistant with retrieval (RAG), JSON mode, safety policies and tool calls, exposed via UI and API.

BASE_URL
/api/practice
Console app
/lab/ai-testing-score-outputs-with-an-llm-judge

Your starter code already declares BASE_URL — call the API relative to it.

Objective

Implement a rubric judge that returns { score, reason } per answer and calibrate it on known good and bad answers.

Your task

  1. 1Write judge(case, response) scoring 0–3: +1 contains expected facts, +1 cites must_cite, +1 refusal matches should_refuse.
  2. 2Score the "standard-shipping" eval case with acme-assistant-1 and assert score === 3.
  3. 3Score the same case with acme-assistant-2 (no citations) and assert score < 3 with a reason mentioning citations.

Acceptance criteria

  • GET /ai/evals returns 200
  • POST /ai/chat returns 200
  • Both models are judged
  • At least 2 assertions pass

LLM evaluation · AI Testing · Evals & regression