LLM evaluation

Test loop and step-limit guards

Expert190 pts~70 min
  • Guardrails
  • Step limits
Practice app · Acme Support Assistant

A deterministic LLM-style support assistant with retrieval (RAG), JSON mode, safety policies and tool calls, exposed via UI and API.

BASE_URL
/api/practice
Console app
/lab/ai-testing-test-loop-and-step-limit-guards

Your starter code already declares BASE_URL — call the API relative to it.

Objective

Assert the assistant's work per request is bounded: few steps, at most one tool call, and a respected token budget.

Your task

  1. 1Send a conversation repeating "Track order 1001" in three user turns with tools: true.
  2. 2Assert steps.length ≤ 5 and tool_calls.length ≤ 1.
  3. 3Send the same with max_tokens: 3 → assert finish_reason === "length" and truncated === true.
  4. 4Assert finish_reason is always one of "stop", "length", "tool_calls".

Acceptance criteria

  • POST /ai/chat returns 200
  • A step or token limit is asserted
  • At least 4 assertions pass

LLM evaluation · AI Testing · Agents & tool use