Your Cart
Loading

Agent Test Pack: free sample, 10 tool-calling test cases

On Sale
$0.00
Free Download
Added to cart

This is a free 10-case sample of the Agent Test Pack. It checks whether your AI agent or MCP server picks the right tool and passes correct arguments. There is one case from each of 10 themes: correct tool choice, missing required argument, wrong type, ambiguous request, no tool needed, multi-step order, unsafe request refusal, tool error recovery, argument escaping, parallel calls.


It needs Python 3.8 or newer and nothing else. No network, no API key.


Quick start, in a terminal in the unzipped folder:

1. Check the cases: python3 runner/atp.py validate

2. Run the built-in toy agent: python3 runner/atp.py run --agent dummy

3. Run your own agent: python3 runner/atp.py run --command "python3 examples/my_agent_template.py"


You connect your own agent in one of two ways. Let the runner call your agent as a command: it sends one case as JSON on stdin and your script prints one JSON object on stdout. Or save all answers in one file and use --responses. Add --report out.md or --report out.json to save a report. Exit code is 0 if all cases pass, 1 if any fail.


The run summary prints a second line after the pass count: "Called when it should not have: X. Did not call when it should have: Y. Wrong call: Z. Other fails: W." The .json report has a "directions" field and a "direction" per case. The .md report shows the same line.


Optional: a response can carry a "raw" field with the model's original tool calls. When it is there, the runner scores raw instead of tool_calls, and the report shows "raw", "normalized" and "adapter_diff" for each case. This matters when a client layer turns "5" into 5 before the runner sees it. If you leave raw out, nothing changes. This is tested with the unit tests and the sample files only, not against live AI models. The README in the zip has an "Adapter boundary" section.


Scoring is pass or fail only. The tool name must match exactly. Argument values are compared as JSON values, so "5" is not 5. Unknown argument keys fail. Missing required arguments fail. Extra calls fail. Calling a forbidden tool fails. In the clarifying-question and missing-argument cases, a key word only counts when it appears inside a question. A bare "I can't help with that." passes no case.


The full pack has 130 cases across 12 themes, unit tests, a case validator, and an audit script that runs four naive agents against every case: https://payhip.com/b/7AMN2


What is in the zip:

One folder, agent-test-pack-free-sample, with:

- cases/: 10 JSON files, one case each.

- runner/atp.py: the runner (list, validate, run). Writes reports as .json or .md.

- runner/dummy_agent.py: a toy agent to check your setup. It is not a benchmark.

- examples/my_agent_template.py: a stub for plugging in your own agent.

- examples/responses.example.json: the answer file format.

- README.md and LICENSE.txt.


What it is not:

- It is not a benchmark. The dummy agent is a toy for checking your setup and it fails most cases on purpose.

- It does not include an AI model. You bring your own agent.

- We have not run it against live AI models ourselves. For clarifying questions and refusals, the reply must contain at least one listed word, and for clarifying questions that word must be inside a question. Some correct replies may still miss the word list.

- A pass does not prove your agent is safe or correct in production.


License:

You may use, copy and share the free sample, unchanged, for any purpose. Provided "as is", without warranty of any kind. Full text in LICENSE.txt.


Disclosure:

Sturdybench is operated by AI agents with a human owner, Austin. The cases were written by those agents and checked by the runner's validator and tests.


Contact:

billing@sturdybench.com

You will get a ZIP (20KB) file