LLM evaluation

Track cost and latency per eval

Hard120 pts~45 min
  • Token usage
  • Latency
  • Cost tracking
Practice app · Acme Support Assistant

A deterministic LLM-style support assistant with retrieval (RAG), JSON mode, safety policies and tool calls, exposed via UI and API.

BASE_URL
/api/practice
Console app
/lab/ai-testing-track-cost-and-latency-per-eval

Your starter code already declares BASE_URL — call the API relative to it.

Objective

Record tokens and latency for each eval case and assert budgets.

Your task

  1. 1Run each case from GET /ai/evals and record usage.total_tokens and latency_ms.
  2. 2Assert total_tokens === prompt_tokens + completion_tokens for every case.
  3. 3Assert every latency_ms < 2000 and the average total_tokens < 200.
  4. 4Print a table of id, tokens, latency.

Acceptance criteria

  • GET /ai/evals returns 200
  • POST /ai/chat returns 200
  • At least 3 assertions pass

LLM evaluation · AI Testing · Evals & regression