
Building a Frontier Legal Evaluation in Partnership with Legora
Sam J.
Spencer M.AfterQuery’s research team worked with Legora to build a frontier legal benchmark. The result was the Benchmark for Agentic Reasoning (BAR), which tests whether agents can complete end-to-end legal work inside Legora’s custom harness.
BAR includes 5,161 cases across 28 practice areas and 11,075 source documents. Legora also published one expert-validated synthetic case developed jointly with AfterQuery, giving researchers a complete view of how the private benchmark’s tasks and scoring work.
Defining the Benchmark
AfterQuery and Legora began by defining what the evaluation needed to reveal to Legora’s team. The teams selected the legal workflows, set the level of realism, defined what agents would see and produce, and established what qualified as a strong answer.
With matters and source documents created alongside participating law firms, the teams turned the underlying legal work into repeatable agent tasks. Each task specified the evidence available to the model, the work product it needed to deliver, and the standards used to score it. This kept BAR grounded in realistic matters while making it consistent enough to run across thousands of cases.
Building the Evaluation Together
AfterQuery and Legora worked closely throughout task design and validation. AfterQuery ran the tasks and shared failure cases with Legora’s team, whose feedback helped distinguish meaningful legal errors from reasonable differences in judgment. The resulting findings shaped the source materials, expected outputs, rubrics, and judges until both teams were confident that the benchmark reflected the work Legora expected its agents to perform.
That process gave Legora a benchmark its research and product teams could trust. It also gave both teams a shared way to review failures and decide what to improve.
Using BAR to improve the product
BAR now helps Legora tell whether a failure comes from the model or the harness and decide what to improve next. In one month, harness changes informed by BAR increased output quality by 5% across the same production models.
To learn more about how AfterQuery can build a custom evaluation set for your business, contact research@afterquery.com.
AfterQuery is an applied research lab curating data solutions to accelerate foundation model development.



