mechub sovereign network-security automation ← all projects
tool early MIT

mechubbench

Agent tool-call benchmarking

Tool-call benchmark corpus and runner for network-automation agents: runs scenarios through local models and scores the emitted tool-call sequences.

About

mechubbench is a benchmark harness for evaluating LLM tool-calling accuracy on network firewall automation tasks. It runs scenarios against local models via an OpenAI-compatible endpoint and scores the emitted tool-call sequences.

Scenarios are YAML files describing firewall-automation tasks, such as auditing trust to untrust policies and staging a fix without applying it.

Features

  • Scenario files list expected tool calls and forbidden calls
  • Scoring checks that all expected calls are present and ordered, with no forbidden call
  • bench run drives scenarios against a model and writes a manifest
  • bench lint validates scenario files

Quick start

Install

uv pip install -e .

Run

bench run --model llama3.2:3b --scenarios scenarios/ --out manifest.json

Full instructions in the README ↗