LLM evaluation

Compute pass rate across a dataset

Medium70 pts~25 min
  • Pass rate
  • Eval harness
Practice app · Acme Support Assistant

A deterministic LLM-style support assistant with retrieval (RAG), JSON mode, safety policies and tool calls, exposed via UI and API.

BASE_URL
/api/practice
Console app
/lab/ai-testing-compute-pass-rate-across-a-dataset

Your starter code already declares BASE_URL — call the API relative to it.

Objective

Run every golden case and compute the pass rate for acme-assistant-1.

Your task

  1. 1GET BASE_URL + "/ai/evals".
  2. 2For each case send the prompt at temperature 0; pass = all expected_contains present AND must_cite cited (if set) AND refused === should_refuse (if set).
  3. 3Compute passRate = passed / total and assert it is ≥ 0.9; print failing ids.

Acceptance criteria

  • GET /ai/evals returns 200
  • POST /ai/chat returns 200
  • The dataset is iterated
  • At least 1 assertion passes

LLM evaluation · AI Testing · Evals & regression