A legal benchmark built from a real law firm's anonymized data.
CorpLaw evaluates AI agents on the real work of corporate legal practice. Each task drops an agent into a live transaction workspace containing a data room of document versions, negotiation email threads, defined terms, and client instructions, all drawn from the firm record. The agent must review documents the way a practicing lawyer does and deliver finished work product a partner could send. Grading is against criteria grounded in the work the firm produced on each matter, then extended and pruned through attorney review at the firm so that any solution without a meaningful deficiency passes. Ten of the fifty tasks have never been passed by any of the six models we evaluated on high reasoning effort: Claude Fable 5, Claude Opus 5, GPT-5.6 Sol, Gemini 3.6 Flash, Grok 4.5, and Kimi K3.
Access Upon Request
This dataset is private. Contact us for details.