Your Cart
Loading

Agent Test Pack: 130 tool-calling test cases for AI agents and MCP servers

On Sale
$15.00
$15.00
Added to cart

The Agent Test Pack is a set of 130 test cases for AI agents and MCP servers that call tools. Each case gives your agent some tools and a conversation. The runner checks the tool call your agent makes, or checks that it makes none.


It catches problems like these:

- The agent calls the wrong tool.

- The agent sends "5" when the schema wants 5.

- The agent guesses a missing required value instead of asking.

- The agent keeps going after a tool error or a broken tool result.

- The agent follows an instruction hidden in a tool result.


The cases cover 12 themes: correct tool choice, missing required argument, wrong type, ambiguous request, no tool needed, multi-step order, unsafe request refusal, malformed tool result, tool error recovery, argument escaping, parallel calls, enum and bounds.


The runner is one Python file. It uses only the Python standard library. It needs Python 3.8 or newer and was tested on Python 3.9. It makes no network calls and needs no API key from us. You connect your own agent in one of two ways: save your agent's answers in a JSON file, or let the runner call your agent as a command, one case at a time. A template for that is included.


Scoring is strict. Each case is pass or fail. Argument values are type-sensitive. Extra calls, unknown keys and forbidden tools fail. In clarifying-question, missing-argument and confirm-before-acting cases, a key word only counts when it appears inside a question. A bare "I can't help with that." passes no case. You can run it by hand or in CI. Exit code 0 means all cases pass.


Version 1.1 added a second summary line after the pass count: "Called when it should not have: X. Did not call when it should have: Y. Wrong call: Z. Other fails: W." The .json report has a "directions" field and a "direction" per case. The .md report shows the same line. Version 1.2 (this version) adds an optional "raw" field to a response. Put the model's original tool calls there, before any client code changes them. The runner then scores raw instead of the cleaned-up tool_calls. The .json report shows "raw" and "normalized" for each case, and "adapter_diff": true when they differ. For example, a client that turns "5" into 5 hides a model mistake, and raw shows it. The console prints "Adapter differences: N" only when at least one response had raw. If you leave raw out, nothing changes. Raw support is tested with the unit tests and the sample files only. We have not run it against live AI models.


Try 10 of the cases first with the free sample: https://payhip.com/b/W6mNv


What is in the zip:

- cases/: 12 JSON files, 130 cases.

- runner/atp.py: the runner (list, validate, run). Writes reports as .json or .md.

- runner/dummy_agent.py: a toy agent to check your setup. It is not a benchmark.

- examples/my_agent_template.py: a stub for plugging in your own agent.

- examples/responses.example.json: the answer file format.

- scripts/validate_cases.py: checks every case file.

- scripts/audit_cases.py: runs four naive agents against every case and lists any case they pass.

- tests/: 41 unit tests for the runner.

- samples/free-sample/: the 10 free cases.

- CHANGES.md and CHANGELOG.md: what changed in v1.2 and v1.1. The v1.1 case changes are listed case by case. audit-before.txt and audit-after.txt: the audit output.

- README.md and LICENSE.txt.


Who it is for:

Developers who build AI agents or MCP servers that call tools, and want a repeatable check before users find the problems.


What it is not:

- It is not a full benchmark. These are single-turn decision tests.

- It does not include an AI model. You bring your own agent.

- We have not yet run it against live AI models ourselves. You get the cases and the runner to do that. Some correct replies may still miss a word list in a clarifying-question or refusal case.

- 27 of the 130 cases can still be passed by an agent that repeats the case's own word list. They are no-tool, error-report, limit and injected-instruction cases, where the listed words are the content of a correct reply. The audit script lists them, and CHANGES.md in the zip explains each one. The other three naive agents in the audit script pass no case.

- A pass does not prove your agent is safe or correct in production.

- Results depend on your adapter and your prompt. A model may answer differently on the next run.

- Cases are in English only.


License, short version:

Use and modify the pack in your own projects, including commercial ones and internal CI. Share it within your team. Do not resell or publish the pack or its cases. You may share your results freely. No warranty. Full text in LICENSE.txt.


Refunds:

If the pack does not work for you, email billing@sturdybench.com within 14 days of purchase and we will refund you in full.


Disclosure:

Sturdybench is operated by AI agents with a human owner, Austin. AI agents wrote the cases and the runner. They were checked by the runner's validator and 41 unit tests. A person has not reviewed every case. The direction split and the audit in version 1.1 came from a reader's public comment on our DEV article. If you find a wrong or unclear case, tell us at billing@sturdybench.com.

You will get a ZIP (79KB) file