LLM evaluation

Test recovery from a failed tool call

Hard120 pts~45 min
  • Error recovery
  • Tool failures
Practice app · Acme Support Assistant

A deterministic LLM-style support assistant with retrieval (RAG), JSON mode, safety policies and tool calls, exposed via UI and API.

BASE_URL
/api/practice
Console app
/lab/ai-testing-test-recovery-from-a-failed-tool-call

Your starter code already declares BASE_URL — call the API relative to it.

Objective

Make the tool fail to find an order and assert the assistant reports it instead of inventing a status.

Your task

  1. 1Ask "Track order 99999" with tools: true.
  2. 2Assert tool_calls[0].name is get_order_status and its result does not contain a status.
  3. 3Assert output_text contains none of pending, paid, shipped, delivered, cancelled, and refused is false.
  4. 4Ask "Track order 1001" with tools: false → assert tool_calls is empty and no status is claimed.

Acceptance criteria

  • POST /ai/chat returns 200
  • Tools are enabled
  • At least 3 assertions pass

LLM evaluation · AI Testing · Agents & tool use