LLM evaluation

Gate a deployment on eval scores in CI

Expert190 pts~70 min
  • Quality gates
  • CI/CD for LLMs
Practice app · Acme Support Assistant

A deterministic LLM-style support assistant with retrieval (RAG), JSON mode, safety policies and tool calls, exposed via UI and API.

BASE_URL
/api/practice
Console app
/lab/ai-testing-gate-a-deployment-on-eval-scores-in-ci

Your starter code already declares BASE_URL — call the API relative to it.

Objective

Write a deployment gate that blocks a candidate model whose eval pass rate is below the threshold.

Your task

  1. 1Implement gate(model, threshold = 0.95) that runs the full eval set and returns { passRate, allowed }.
  2. 2Assert gate("acme-assistant-1").allowed === true.
  3. 3Assert gate("acme-assistant-2").allowed === false and print the failing case ids.

Acceptance criteria

  • GET /ai/evals returns 200
  • POST /ai/chat returns 200
  • Both models are gated
  • At least 2 assertions pass

LLM evaluation · AI Testing · Evals & regression