Measure answer faithfulness
Hard120 pts~45 min
- Faithfulness
- RAG metrics
Practice app · Acme Support Assistant
A deterministic LLM-style support assistant with retrieval (RAG), JSON mode, safety policies and tool calls, exposed via UI and API.
Your starter code already declares BASE_URL — call the API relative to it.
Objective
Compute a faithfulness score over several questions: the share of answers whose key fact is in a cited document.
Your task
- 1Load the KB with GET BASE_URL + "/ai/kb".
- 2For refunds (“30 days”), warranty (“12-month”), price match (“14 days”) and gift cards (“non-refundable”) ask the question.
- 3Mark an answer faithful when the fact is in output_text AND in the content of a cited doc.
- 4Assert faithfulness === 1.0 and print the per-case results.
Acceptance criteria
- GET /ai/kb returns 200
- POST /ai/chat returns 200
- At least 2 assertions pass
LLM evaluation · AI Testing · Hallucination & grounding (RAG)