LLM evaluation

Test a multi-step agent trajectory

Hard120 pts~45 min
  • Agent trajectories
  • Step assertions
Practice app · Acme Support Assistant

A deterministic LLM-style support assistant with retrieval (RAG), JSON mode, safety policies and tool calls, exposed via UI and API.

BASE_URL
/api/practice
Console app
/lab/ai-testing-test-a-multi-step-agent-trajectory

Your starter code already declares BASE_URL — call the API relative to it.

Objective

Assert the ordered steps the assistant took for a tool task and for a knowledge task.

Your task

  1. 1Ask "Track order 1001" with tools: true → assert steps types are ["tool", "answer"] in order.
  2. 2Ask "What warranty do your products have?" → assert steps types are ["retrieve", "answer"].
  3. 3Assert the final step is always "answer" and the tool step detail mentions get_order_status.

Acceptance criteria

  • POST /ai/chat returns 200
  • Tools are enabled
  • At least 3 assertions pass

LLM evaluation · AI Testing · Agents & tool use