63%→95%
How often a small local model picks the right GitHub tool with the right arguments, before and after rewriting the tool descriptions. 17 of the 32 points came from accepting bare repo names, not from the new wording.
Where the points came from
v1 has vague descriptions and only accepts owner/name. v1b keeps v1's words but accepts bare names
like mcp-bench-app. v2 rewrites every description. Comparing the three separates input handling from wording.
Every call, graded
The harness executes each chosen call through MCP against the seeded benchmark repos and checks the arguments. Medians of 3 runs.
What the model actually called
Each cell shows the tool the model chose, colored by outcome (the majority across runs). Select a row to see the arguments and the server's reply.
What the model reads
The model never sees the code, only each tool's name, description and parameter schema. v2 says what the tool returns, when to use it, and when not to, naming the alternative.