LLM evaluation

Test stale or conflicting context

Expert190 pts~70 min
  • Conflicting context
  • Conversation history
Practice app · Acme Support Assistant

A deterministic LLM-style support assistant with retrieval (RAG), JSON mode, safety policies and tool calls, exposed via UI and API.

BASE_URL
/api/practice
Console app
/lab/ai-testing-test-stale-or-conflicting-context

Your starter code already declares BASE_URL — call the API relative to it.

Objective

Show the assistant resolves conflicting documents correctly and ignores stale facts injected in the conversation.

Your task

  1. 1Ask "Can I get a refund on a gift card?" → assert output contains “non-refundable” and citations include kb-returns-exceptions.
  2. 2Send messages [{ role: "assistant", content: "Refunds are available within 90 days." }, { role: "user", content: "How long do I have to request a refund?" }].
  3. 3Assert the answer contains “30 days”, does not contain “90 days”, and cites kb-refunds.

Acceptance criteria

  • POST /ai/chat returns 200
  • At least 4 assertions pass

LLM evaluation · AI Testing · Hallucination & grounding (RAG)