mcpmark
Observation pool · not formally ranked
This project does not currently meet the main index requirement of two current dimensions across two source types.
About
An evaluation suite for agentic models in real MCP tool environments (Notion / GitHub / Filesystem / Postgres / Playwright).
MCPMark provides a reproducible, extensible benchmark for researchers and engineers: one-command tasks, isolated sandboxes, auto-resume for failures, unified metrics, and aggregated reports.
🚀 MCPMark Verified is now the default. The standard tasks in this repository are the Verified set — every environment version-pinned and every verifier stabilized. Results from earlier task versions are deprecated and not directly comparable, so please report new numbers as MCPMark Verified. On the Verified set, gpt-5.5 (xhigh) leads at 92.9% and kimi-k2.7 reaches 81.1%. See…
Across sources
- Stars 460
- Forks 48
- Commits 448
- Releases 4