Harvey LAB-AA v1.1 introduces hallucination checks and new metrics
Summary
Today, Artificial Analysis announced the release of Harvey LAB-AA v1.1, an updated scoring methodology for their Legal Agent Benchmark (LAB) in collaboration with Harvey. This version introduces a Hallucination-Gated All-Pass Rate, which now requires that responses not only meet rubric criteria but also contain no material hallucinations—errors that might mislead on substantive points. Harvey's human preference studies indicated that hallucinations significantly influence lawyers' choices between AI outputs. The update reflects an effort to improve the evaluation of legal AI outputs, ensuring that only those devoid of misleading errors receive credit, thereby enhancing the reliability of the assessment process for legal tasks.
Analysis
Harvey: Harvey is a legal technology company focused on building AI tools and benchmarks to support lawyers in complex, long-horizon legal work. It partnered on the v1.1 update to its Legal Agent Benchmark (LAB-AA), adding a hallucination check and multi-judge rubric evaluation. The company maintains a private dataset of realistic legal tasks across practice areas and collaborates with model developers on evaluation methodology. OpenAI: OpenAI develops advanced AI models including the GPT-6 family. Its GPT-6 Astra (max) model placed third on the Hallucination-Gated All-Pass Rate in Harvey LAB-AA v1.1 and showed strong grounding with very low material hallucination rates. The company’s models were evaluated across the updated legal benchmark tasks. Kimi K3: Kimi K3 is an AI model from Moonshot. It was evaluated on the Harvey LAB-AA v1.1 benchmark and showed a high criterion pass rate but elevated material hallucination frequency. The model was referenced in related human preference and GRM studies by Harvey. AIatMeta: AIatMeta develops AI models including the Muse Spark series. Its Muse Spark 1.3 (max) model ranked second on the Hallucination-Gated All-Pass Rate in the updated Harvey LAB-AA v1.1 leaderboard after the hallucination filter was applied. The company participates in public evaluations of model performance on legal agent tasks. Grok 4.7: Grok 4.7 is a high-capability AI model developed by SpaceXAI. It led the Harvey LAB-AA v1.1 benchmark on the primary Hallucination-Gated All-Pass Rate metric. The model was tested on a private set of 120 legal tasks in collaboration with Harvey. SpaceXAI: SpaceXAI develops frontier AI models including the Grok family. Its Grok 4.7 (xhigh) model achieved the highest score on the new Hallucination-Gated All-Pass Rate metric in the Harvey LAB-AA v1.1 benchmark. The company was thanked for its contributions to the benchmark collaboration. GPT-6 Sol: GPT-6 Sol is a model in OpenAI’s GPT-6 family. It was selected as the production hallucination checker for Harvey LAB-AA v1.1 and identified more material hallucinations than other checkers in the 20-task subset. The model also participated in the three-judge rubric panel. GPT-6 Luna: GPT-6 Luna is a GPT-6 family model from OpenAI. It appeared among the Pareto frontier models for score versus cost on the Harvey LAB-AA v1.1 benchmark. The model was tested alongside other GPT-6 variants on the legal agent tasks. nikogrupen: Nikogrupen is a member of the Harvey team who contributed to the development and collaboration on the Harvey LAB-AA benchmark. He was specifically thanked for work on the v1.1 methodology updates. GPT-6 Astra: GPT-6 Astra is a model in OpenAI’s GPT-6 family. It achieved third place on the Hallucination-Gated All-Pass Rate in Harvey LAB-AA v1.1 and demonstrated the lowest material hallucination rate among tested models. The model retained nearly all of its rubric passes after the hallucination filter was applied. Julio Pereyra: Julio Pereyra is a Harvey team member recognized for contributions to the Harvey LAB and the v1.1 benchmark collaboration. His work supported the partnership with model providers on the updated evaluation framework. Muse Spark 1.3: Muse Spark 1.3 is an AI model developed by AIatMeta. It ranked second on the updated Harvey LAB-AA v1.1 leaderboard after the addition of a hallucination gate to scoring. The model performed strongly on rubric criteria but saw a significant drop once material hallucinations were penalized. Claude Opus 5.5: Claude Opus 5.5 is an AI model developed by Anthropic. It served as one of three judges in the updated rubric panel for Harvey LAB-AA v1.1 and was evaluated as a hallucination checker on a 20-task subset. The model contributed to the averaged scoring used to mitigate bias. Gemini 3.8 Flash: Gemini 3.8 Flash is a model from Google. It was included in the six-model hallucination checker comparison on Harvey LAB-AA tasks. The model identified fewer material hallucinations than GPT-6 Sol or Grok 4.7 in the tested subset. Claude Sonnet 5.5: Claude Sonnet 5.5 is an AI model from Anthropic. It participated in the Harvey LAB-AA v1.1 evaluation both as a model under test and as a hallucination checker. The model showed lower detection rates for material hallucinations compared with some other checkers in the 20-task comparison. Legal AI Focus: Human preference studies by Harvey identified hallucinations as a primary factor lawyers use when choosing between otherwise comprehensive AI outputs. Benchmark Methodology: Harvey updated its Legal Agent Benchmark to penalize material hallucinations that could mislead on substantive legal points, distinguishing them from minor errors. Evaluation Collaboration: The v1.1 update was developed through direct collaboration between Artificial Analysis and Harvey using the company’s private legal task dataset.
Categories
aiai_agentsmachine_learningtech