tool
early
MIT
mechubbench
Agent tool-call benchmarking
Tool-call benchmark corpus and runner for network-automation agents: runs scenarios through local models and scores the emitted tool-call sequences.
About
mechubbench is a benchmark harness for evaluating LLM tool-calling accuracy on network firewall automation tasks. It runs scenarios against local models via an OpenAI-compatible endpoint and scores the emitted tool-call sequences.
Scenarios are YAML files describing firewall-automation tasks, such as auditing trust to untrust policies and staging a fix without applying it.
Features
- Scenario files list expected tool calls and forbidden calls
- Scoring checks that all expected calls are present and ordered, with no forbidden call
- bench run drives scenarios against a model and writes a manifest
- bench lint validates scenario files
Quick start
Install
uv pip install -e .Run
bench run --model llama3.2:3b --scenarios scenarios/ --out manifest.json