Tool interaction — calls, responds, matching rules, and the tool_call evaluator.
A user sees a conversation. A tester sees the full round-trip: the assistant deciding to call a tool, what parameters it passes, what comes back, and what the assistant says about it afterward. ABS makes every step of that round-trip assertable.
A tool call is two Behaviors, not one:
# The assistant calls a tool
- actor: assistant
action: calls
target: Order MCP
with:
orderId: "12345"
# The tool responds
- actor: tool
action: responds
target: Order MCP
content:
status: "shipped"
eta: "2026-07-30"
Two Behaviors, back to back. The first says what was called and with what parameters. The second says what came back. The assistant's follow-up (informs, shows, etc.) comes after.
This separation is deliberate: the tool response payload is assertable on its own (schema validation, exact values), independently of how the assistant later paraphrases it.
with: partial (default) vs. strictwith — partial match (default)The observed tool arguments must contain every key in with, but may contain additional keys:
# ABS says:
with:
orderId: "12345"
# Agent actually called:
{ "orderId": "12345", "currency": "USD", "priority": "high" }
# → MATCH. orderId matches. Extra keys are fine.
This is the default because it makes sessions resilient: the agent can pass contextual parameters (locale, tracing IDs, internal flags) without breaking the spec.
with_only — strict matchThe observed tool arguments must match exactly — no extra keys allowed:
# ABS says:
with_only:
orderId: "12345"
# Agent actually called:
{ "orderId": "12345", "currency": "USD" }
# → NO MATCH. currency is present but not declared.
Use with_only when the exact set of parameters matters — security-critical calls, compliance boundaries, or when an extra parameter would change behavior.
with and with_only are mutually exclusive on the same Behavior. If neither is present, the Behavior only checks that the tool was called (target match) without inspecting parameters.
An agent can call several tools before saying anything to the user. ABS represents this as consecutive calls Behaviors:
- actor: assistant
action: calls
target: Order MCP
with:
orderId: "12345"
- actor: assistant
action: calls
target: Inventory MCP
with:
sku: "SKU-001"
By default, multiple calls must match in the order they appear in the session.
If order doesn't matter — the agent can call them in any sequence — use the tool_call evaluator with ordered: false:
- actor: assistant
action: calls
target: Order MCP
- actor: assistant
action: calls
target: Inventory MCP
evaluations:
- type: tool_call
ordered: false
With ordered: false, both calls must occur, but in either order.
When the session declares a tool responds Behavior, the Runner matches it against the tool result the agent received:
- actor: tool
action: responds
target: Order MCP
content:
status: "shipped"
This checks that:
Order MCP was receivedcontentStructural compatibility means: if content is an object, the observed payload must have at least those keys with those types. Extra keys are fine (same philosophy as with partial matching).
For strict payload matching, use the schema evaluator:
- actor: tool
action: responds
target: Order MCP
content:
status: "shipped"
evaluations:
- type: schema
schema:
type: object
required: [status, eta]
properties:
status: { type: string, enum: [shipped, pending, cancelled] }
eta: { type: string, format: date }
additionalProperties: false
tool_call evaluatorThe tool_call evaluator checks one or more calls Behaviors against what the agent actually invoked. It can be placed on an individual calls Behavior or on the last calls in a group.
evaluations:
- type: tool_call
target: "Order MCP" # Optional — already on the Behavior, but can be explicit
with: # Optional — same
orderId: "{{orderId}}"
ordered: true # Optional — default true
| Field | Default | Description |
|---|---|---|
target | (from Behavior) | Which tool should have been called |
with | (from Behavior) | Parameters that must be present |
ordered | true | If multiple calls in a row, must they be in this order? |
When there's nothing to validate beyond what the Behavior already declares (target + with), the tool_call evaluator is optional — the Runner already matches target and with/with_only as part of step matching. Use the evaluator when you want to assert additional properties like ordered: false.
| Question | Answer |
|---|---|
| Default parameter matching | with = partial (observed must contain all keys; extra keys OK) |
| Strict parameter matching | with_only = exact (no extra keys allowed) |
| Multiple tools, ordering | Default ordered: true. Set ordered: false in tool_call evaluator for independent calls |
| Tool response matching | Structural compatibility by default. Use schema evaluator for strict validation |
| Tool response payload assertion | Use the tool Behavior's own evaluations block |
ABS is protocol-agnostic, but the calls / with / responds pattern maps naturally to MCP (Model Context Protocol):
calls → MCP tool invocationwith → MCP tool parametersresponds → MCP tool resultThe same pattern works for REST APIs, GraphQL, function calling, or plugin systems.