LLM evaluation

Build a golden dataset of cases

Medium70 pts~25 min
  • Golden datasets
  • Evals
Practice app · Acme Support Assistant

A deterministic LLM-style support assistant with retrieval (RAG), JSON mode, safety policies and tool calls, exposed via UI and API.

BASE_URL
/api/practice
Console app
/lab/ai-testing-build-a-golden-dataset-of-cases

Your starter code already declares BASE_URL — call the API relative to it.

Objective

Load the golden dataset, validate its structure and run the first cases against the assistant.

Your task

  1. 1GET BASE_URL + "/ai/evals" and assert it is a non-empty array.
  2. 2Assert every case has id, prompt and a non-empty expected_contains array; ids are unique.
  3. 3Run three cases and assert each output contains all of its expected_contains strings.

Acceptance criteria

  • GET /ai/evals returns 200
  • POST /ai/chat returns 200
  • At least 4 assertions pass

LLM evaluation · AI Testing · Evals & regression