Research Built to Be Used

We build benchmarks, datasets, and infrastructure to advance machine learning research.

We are a research organization building open benchmarks, datasets, and infrastructure. Our work explores machine learning, autonomous systems, and how computational tools can advance scientific research.


Benchmarks

We build benchmarks to measure how AI agents perform on real research tasks.

Datasets

1.1M enriched papers. 129K research repositories. 778K code functions. The raw material for studying how AI systems interact with real scientific work.

Infrastructure

Runtimes for structured agent workloads. Orchestration topologies for recursive improvement loops. The scaffolding to run experiments at scale.

S2ORC CS Enriched: 1.1 Million Computer Science Papers with Structured Metadata

A filtered and LLM-enriched version of Allen AI's S2ORC corpus containing 1.1 million computer science papers with structured extraction...

datasets / scientific-papers / machine learning

Study Failure: AI-driven GPU Kernel Optimization

A retrospective on 131,520 GPU kernel optimization attempts that were invalidated when agents were found to be substituting high-level...

gpu / optimization / machine learning

Learning to Rank Architectures: A Small Model That Guides Neural Architecture Search

A tiny recursive reasoning model trained to rank architectures by predicted performance achieves 8-10x sample efficiency over random...

nas / architecture search / machine learning

ARIA Benchmark: How Much Machine Learning Do AI Models Actually Know?

A suite of five closed-book benchmarks probing the ML knowledge that frontier language models have internalized during training.

agent-evaluation / benchmarks / python

ArXiv Research Code Dataset: 129K Research Repositories

A collection of 4.7 million code files from 129K research repositories linked to arXiv computer science papers.

agent-evaluation / benchmarks / python

ArXivDLInstruct: 778K Research Code Functions for Instruction Tuning

A dataset of 778,152 functions extracted from arXiv-linked research code, each paired with instruction prompts, for training...

agent-evaluation / benchmarks / python

DeltaML-Bench: Evaluating Machine Learning Agents on Real-World Research Repositories

A 48-task benchmark for evaluating autonomous ML experimentation in real research repositories.

agent-evaluation / benchmarks / python

Teaching Models to Bluff: Measuring Deception, Belief, and Coordination in LLM Secret Hitler

We implemented five LLM agents playing the social-deduction game Secret Hitler with structured logging to quantify deception, belief...

ai-research / agi / recursive-improvement

ML Research Benchmark: Can AI Agents Do Real ML Research?

A benchmark suite of 7 competition-level ML challenges for evaluating whether AI agents can perform genuine research iteration beyond...

agent-evaluation / benchmarks / python

Algorithmic Research Group builds tools and infrastructure for research. Benchmarks for evaluating autonomous agents. Datasets for studying machine learning systems. Runtimes for running agent workloads at scale.