Score outputs with an LLM judge
Hard120 pts~45 min
- LLM-as-judge
- Rubrics
Practice app · Acme Support Assistant
A deterministic LLM-style support assistant with retrieval (RAG), JSON mode, safety policies and tool calls, exposed via UI and API.
Your starter code already declares BASE_URL — call the API relative to it.
Objective
Implement a rubric judge that returns { score, reason } per answer and calibrate it on known good and bad answers.
Your task
- 1Write judge(case, response) scoring 0–3: +1 contains expected facts, +1 cites must_cite, +1 refusal matches should_refuse.
- 2Score the "standard-shipping" eval case with acme-assistant-1 and assert score === 3.
- 3Score the same case with acme-assistant-2 (no citations) and assert score < 3 with a reason mentioning citations.
Acceptance criteria
- GET /ai/evals returns 200
- POST /ai/chat returns 200
- Both models are judged
- At least 2 assertions pass
LLM evaluation · AI Testing · Evals & regression