LLM evaluation

Measure answer faithfulness

Hard120 pts~45 min
  • Faithfulness
  • RAG metrics
Practice app · Acme Support Assistant

A deterministic LLM-style support assistant with retrieval (RAG), JSON mode, safety policies and tool calls, exposed via UI and API.

BASE_URL
/api/practice
Console app
/lab/ai-testing-measure-answer-faithfulness

Your starter code already declares BASE_URL — call the API relative to it.

Objective

Compute a faithfulness score over several questions: the share of answers whose key fact is in a cited document.

Your task

  1. 1Load the KB with GET BASE_URL + "/ai/kb".
  2. 2For refunds (“30 days”), warranty (“12-month”), price match (“14 days”) and gift cards (“non-refundable”) ask the question.
  3. 3Mark an answer faithful when the fact is in output_text AND in the content of a cited doc.
  4. 4Assert faithfulness === 1.0 and print the per-case results.

Acceptance criteria

  • GET /ai/kb returns 200
  • POST /ai/chat returns 200
  • At least 2 assertions pass

LLM evaluation · AI Testing · Hallucination & grounding (RAG)