Browse
Collections
Built on real production data, measured against frontier models.
Get in Touch
Pick a Time
Prefer to fill out our interest form?
↗
New
ShelfLife
A digital twin of a real e-commerce company.
200 tasks · measured Aug 13, 2026
Top models · Pass@1
Claude Opus 5
69.0%
GPT-5.6 Sol
55.7%
Claude Fable 5
54.7%
Accounting
Marketing
Finance
Due Diligence & Operational Discipline
Enterprise-Scale Context Awareness
Ethical Conduct
→
Access Upon Request
ShelfLife E-Sim
A stateful simulator of retail operations.
12 tasks · measured Aug 14, 2026
Top models · Mean Reward (pts)
GPT-5.6 Sol
23.0
Claude Fable 5
10.4
Claude Opus 5
6.8
Pricing & Inventory
Retail Operations
Long-Horizon Turn-Based Business Simulation
Multi-Strategy Resource Management
→
Access Upon Request
CorpLaw
A legal benchmark built from a real law firm's anonymized data.
50 tasks · measured Aug 18, 2026
Top models · Pass@1
Claude Opus 5
42.8%
Claude Fable 5
39.5%
GPT-5.6 Sol
36.0%
Legal Writing & Review
Client Counsel
Negotiation
Professional Discipline & Judgement
Legal Issue Severity Assessment
→
Access Upon Request
Long-Horizon SWE
Repo tasks in Go, Python, Rust, C++, TS, and Java. Each task has 100+ steps and a pass rate below 50% for frontier models.
→
Access Upon Request
Cybersecurity
CVE finding, patching, and exploitation across public and private codebases.
→
Access Upon Request
Recursive Self-Improvement
Model training, data curation, inference optimization and harness engineering.
→
Access Upon Request
Terminal Agents
Execution environment setup, enterprise API and tool use, data cleaning and anonymization.
→
Request a New Dataset
We make bespoke evaluations for your model. Tell us what to measure next.
→