aiagent.club
中文
GitHub

mcpmark

Observation pool · not formally ranked

This project does not currently meet the main index requirement of two current dimensions across two source types.

Not ranked
Observed adoptionMissing
MomentumCurrent · 2026-09-12
0
AttentionCurrent · 2026-09-12
0
Signal confidence Based on how many independent score dimensions currently have data.
Medium2/3 · 1 source types

Methodology v2.0 · snapshot 2026-09-12 · stale after 2 days

About

An evaluation suite for agentic models in real MCP tool environments (Notion / GitHub / Filesystem / Postgres / Playwright).

MCPMark provides a reproducible, extensible benchmark for researchers and engineers: one-command tasks, isolated sandboxes, auto-resume for failures, unified metrics, and aggregated reports.

🚀 MCPMark Verified is now the default. The standard tasks in this repository are the Verified set — every environment version-pinned and every verifier stabilized. Results from earlier task versions are deprecated and not directly comparable, so please report new numbers as MCPMark Verified. On the Verified set, gpt-5.5 (xhigh) leads at 92.9% and kimi-k2.7 reaches 81.1%. See…

Across sources

460 Stars
  • Stars 460
  • Forks 48
  • Commits 448
  • Releases 4