Arbor
Observation pool · not formally ranked
This project does not currently meet the main index requirement of two current dimensions across two source types.
About
The autonomous research agent that beats Claude Code and Codex by 2.5× on the same compute budget
Give Arbor a benchmark and a goal. It proposes hypotheses, edits code, runs real experiments, and keeps only the gains that survive held-out data — growing a hypothesis tree instead of forgetting what failed.
▶️ Try it in 30 seconds — no API key, no config: > > > Or watch it right now in your browser — nothing to install: ▶️ Live Demo.
and a metric to measure, from model training to harness engineering to data synthesis.
results, failure modes, and distilled insights in the Idea Tree and propagates them upward, so later ideas start smarter instead of scrolling off.
held-out test split, and…
Across sources
- Stars 1.1k
- Forks 126
- Commits 234
- Releases 4