github-mcp · tool-selection benchmark

63%→95%

How often a small local model picks the right GitHub tool with the right arguments, before and after rewriting the tool descriptions. 17 of the 32 points came from accepting bare repo names, not from the new wording.

The ablation

Where the points came from

v1 has vague descriptions and only accepts owner/name. v1b keeps v1's words but accepts bare names like mcp-bench-app. v2 rewrites every description. Comparing the three separates input handling from wording.

Outcomes

Every call, graded

The harness executes each chosen call through MCP against the seeded benchmark repos and checks the arguments. Medians of 3 runs.

Question by question

What the model actually called

Each cell shows the tool the model chose, colored by outcome (the majority across runs). Select a row to see the arguments and the server's reply.

The rewrite

What the model reads

The model never sees the code, only each tool's name, description and parameter schema. v2 says what the tool returns, when to use it, and when not to, naming the alternative.

Multi-step

Agent mode