Tests

Tests are re-runnable prompts. Instead of creating a new agent from scratch every time, you define a test once and run it repeatedly. Each run creates a new agent linked back to the test, giving you a history of results for the same test over time.

Overview

A test is an abstraction on top of the agent API. Where creating an agent is a one-off operation (one prompt, one run, one result), a test lets you define a prompt once and run it as many times as you want. Each run creates a new agent that is linked back to the test, so you build up a history of results for the same test over time.

This is the core mechanism for tracking the health of your application. If you have a test called "checkout flow" and you run it after every deploy, you get a timeline of pass/fail results that tells you exactly when something broke and when it was fixed. The project health board aggregates this across all tests in a project.

When you run a test, the system resolves the role (cascading from the run parameters, to the test defaults, to the project defaults), copies any attached mailboxes onto the new agent, renders the prompt, and queues the agent for execution.

Prompt template model

A test's prompt_template is the instruction set the agent executes. When a test is run as part of a test plan, everything recorded by ancestor steps is provided to the agent automatically — refer to those values in plain language (“log in with the credentials from the register step”). The only template syntax is for plan-level variables:

  • {{ plan.variables.<name> }} — replaced with a plan-level variable supplied when the run starts.

The starting URL is no longer a separate field. Inline it as literal text in the prompt body, typically as the opening “Navigate to …” line. The agent reads that and calls browser_navigate itself.

A minimal example:

Navigate to {{ plan.variables.env_url }} and log in with the credentials recorded by the register step. Open the order created by the checkout step and verify its status is “confirmed”.

When this test runs via POST /api/v1/tests/{id}/run (outside a plan), plan.variables references render as empty strings, gracefully degrading.

Success determination

Whether a test run succeeded or failed is determined by a two-step process. First, the agent itself reports its own assessment when it finishes (success or failure with a summary). Second, a server-side verification step independently analyzes the full conversation trace and produces its own verdict.

If the agent and the server-side verification disagree, the server-side verdict wins. This means an agent cannot incorrectly report success if the evidence in the conversation shows otherwise, and vice versa. The final outcome is what you see in the API response and on the health board.

Role resolution

When a test is run, the role is resolved using a cascading fallback. The system first checks if a role was specified in the run parameters. If not, it falls back to the default set on the test itself. If the test has no default, it uses the project-level default.

This lets you set sensible defaults on the test or project and override them on individual runs when needed. For example, you might have a test that normally runs with a "Functional QA" role but occasionally run it with a "Power User" role to test a different interaction style.

POST /api/v1/tests

Create a new test

Parameters

ParameterTypeInRequiredDescription
project_iduuidbodyYesID of the project this test belongs to
namestringbodyYesTest name
descriptionstringbodyNoTest description
prompt_templatestringbodyYesPrompt body containing the instructions the agent executes. Inline the starting URL as literal text — the test no longer has a separate entry_url field. See the section overview for the complete template data model.
sourcestringbodyNoOrigin of the test (e.g. 'discovery', 'manual')
statusstringbodyNoLifecycle status: draft (not eligible for regression/schedules), active (eligible), or archived. Defaults to active for manually created tests, draft for generated ones. (default: active)
importancestringbodyNoTest importance level (e.g. 'critical', 'high', 'medium', 'low')
file_pathsstring[]bodyNoPaths of tenant files to copy into the test run workspace when running this test
mailbox_namesstring[]bodyNoOptional list of mailbox names to attach to the test. Every run of the test gains read access to these inboxes. Can be overridden at run time.

Status Codes

CodeDescription
201Test created
400Validation error
401Unauthorized

Response Body

{
  "id": "770e8400-e29b-41d4-a716-446655440000",
  "tenant_id": "550e8400-e29b-41d4-a716-446655440000",
  "project_id": "660e8400-e29b-41d4-a716-446655440000",
  "name": "Checkout Flow",
  "description": "",
  "prompt_template": "Navigate to https://example.com and complete checkout",
  "status": "active",
  "file_paths": [],
  "created_at": "2025-01-15T10:30:00Z",
  "updated_at": "2025-01-15T10:30:00Z"
}
POST /api/v1/tests
cURL
Response
GET /api/v1/tests

List tests

Parameters

ParameterTypeInRequiredDescription
limitintegerqueryNoNumber of results to return (default: 20)
cursoruuidqueryNoCursor for pagination
project_iduuidqueryNoFilter by project ID (alias: project)
statusstringqueryNoFilter by lifecycle status: draft, active, or archived

Status Codes

CodeDescription
200OK
401Unauthorized

Response Body

{
  "tests": [
    {
      "id": "770e8400-e29b-41d4-a716-446655440000",
      "tenant_id": "550e8400-e29b-41d4-a716-446655440000",
      "project_id": "660e8400-e29b-41d4-a716-446655440000",
      "name": "Checkout Flow",
      "description": "",
      "prompt_template": "Navigate to https://example.com and complete checkout",
      "status": "active",
      "file_paths": [],
      "created_at": "2025-01-15T10:30:00Z",
      "updated_at": "2025-01-15T10:30:00Z"
    }
  ],
  "next_cursor": "880e8400-e29b-41d4-a716-446655440000"
}
GET /api/v1/tests
cURL
Response
GET /api/v1/tests/{id}

Get test by ID

Parameters

ParameterTypeInRequiredDescription
iduuidpathYesTest ID

Status Codes

CodeDescription
200OK
400Invalid UUID
401Unauthorized
404Test not found

Response Body

{
  "id": "770e8400-e29b-41d4-a716-446655440000",
  "tenant_id": "550e8400-e29b-41d4-a716-446655440000",
  "project_id": "660e8400-e29b-41d4-a716-446655440000",
  "name": "Checkout Flow",
  "description": "",
  "prompt_template": "Navigate to https://example.com and complete checkout",
  "status": "active",
  "file_paths": [],
  "created_at": "2025-01-15T10:30:00Z",
  "updated_at": "2025-01-15T10:30:00Z"
}
GET /api/v1/tests/{id}
cURL
Response
PATCH /api/v1/tests/{id}

Update a test

Parameters

ParameterTypeInRequiredDescription
iduuidpathYesTest ID
namestringbodyNoTest name
descriptionstringbodyNoTest description
prompt_templatestringbodyNoPrompt body containing the instructions the agent executes. See the section overview.
sourcestringbodyNoOrigin of the test
statusstringbodyNoLifecycle status: draft, active, or archived
importancestringbodyNoTest importance level
file_pathsstring[]bodyNoPaths of tenant files to copy into the test run workspace when running this test
mailbox_namesstring[]bodyNoReplace the test's attached mailbox set with this list. Omit the field to leave attachments unchanged; pass an empty array to detach all mailboxes.

Status Codes

CodeDescription
200Test updated
400Validation error
401Unauthorized
404Test not found

Response Body

{
  "id": "770e8400-e29b-41d4-a716-446655440000",
  "tenant_id": "550e8400-e29b-41d4-a716-446655440000",
  "project_id": "660e8400-e29b-41d4-a716-446655440000",
  "name": "Checkout Flow",
  "description": "",
  "prompt_template": "Navigate to https://example.com and complete checkout",
  "status": "active",
  "file_paths": [],
  "created_at": "2025-01-15T10:30:00Z",
  "updated_at": "2025-01-15T10:30:00Z"
}
PATCH /api/v1/tests/{id}
cURL
Response
DELETE /api/v1/tests/{id}

Delete a test

Parameters

ParameterTypeInRequiredDescription
iduuidpathYesTest ID

Status Codes

CodeDescription
204Test deleted
400Invalid UUID
401Unauthorized
404Test not found
DELETE /api/v1/tests/{id}
cURL
Response
POST /api/v1/tests/{id}/run

Run a test

Creates a new test run from the test, inheriting the test's stored device, browser, and mobile app configuration. Files attached to the test are automatically copied into the test run workspace.

Parameters

ParameterTypeInRequiredDescription
iduuidpathYesTest ID
role_iduuidbodyNoRole for prompt injection context. Overrides the test's default role for this run.
mailbox_namesstring[]bodyNoOverride the test's attached mailboxes for this run. When omitted, the test's stored attachments are inherited.
modelstringbodyNoLLM model override for this run
thinking_levelstringbodyNoOverride the executor model's reasoning depth for this run: MINIMAL, LOW, MEDIUM, or HIGH. Omit to use the model default.
auditor_thinking_levelstringbodyNoReasoning depth for the auditor that reviews this run (MINIMAL, LOW, MEDIUM, HIGH).
max_iterationsintegerbodyNoMax iterations override for this run

Status Codes

CodeDescription
201Agent created from test
400Validation error or template rendering failed
401Unauthorized
404Test not found

Response Body

{
  "id": "550e8400-e29b-41d4-a716-446655440000",
  "tenant_id": "660e8400-e29b-41d4-a716-446655440000",
  "project_id": "770e8400-e29b-41d4-a716-446655440000",
  "status": "pending",
  "prompt": "Navigate to https://example.com and complete checkout",
  "model": "gemini-3.5-flash",
  "browser_type": "chrome",
  "is_discovery": false,
  "messages": [],
  "iteration": 0,
  "max_iterations": 200,
  "tokens_used": 0,
  "test_id": "990e8400-e29b-41d4-a716-446655440000",
  "source": "api",
  "created_at": "2025-01-15T10:30:00Z",
  "updated_at": "2025-01-15T10:30:00Z"
}
POST /api/v1/tests/{id}/run
cURL
Response