Sakana AI’s LLM Peer Review System Catches 73% of Core-Claim Errors


Sakana AI has published Beyond Imitation, a TMLR research paper on LLM-assisted peer review built around error detection. Most AI reviewers are graded on how closely they copy human reviews. This work asks a harder question: can an AI reviewer find a planted mistake? The research team ships two pieces: a Contradiction Benchmark and a Multi-Layered Review (MLR) system. For developers building research agents, the lesson is practical. Both system design and model choice move error detection.

TL;DR

  • Size: 1,164 inserted contradictions across 257 papers from 5 venues. MLR reads up to 10 pages of main text.
  • Runs on: Off-the-shelf API models (Claude Sonnet 4, Claude Haiku 3.5). No GPU, no fine-tuning. About $0.47 per review.
  • Performance: Highest error detection of all 4 systems tested, with human-aligned scores.
  • Best: Caught 73.43% of core-claim errors with 4 reviews, versus 14.81% for the best baseline.
  • Worst: Only 16.11% exact matches on real retracted arXiv papers.
  • Bottom line:
    • Best: reads before judging, and finds far more serious errors.
    • Worst: still falls for hidden prompt injection.

What is Multi-Layered Review?

Multi-Layered Review is an agentic AI review system from Sakana AI that understands a research paper before critiquing it. It uses 3 agents on off-the-shelf Claude models:

  • Appendix Agent (Claude Haiku 3.5): summarizes experiments and implementation details from the appendix.
  • Literature Review Agent (Claude Sonnet 4): uses web search to place the paper in prior work. It is optional.
  • Review Agent (Claude Sonnet 4): runs a 3-pass prompt chain inspired by Keshav’s Three-Pass Approach.

Pass 1 writes a high-level outline. Pass 2 reads in detail and flags weaknesses, assumptions and gaps. Pass 3 merges all agent outputs into Strengths, Weaknesses, Questions, Recommendation, Score and a To-Do list. The PDF is passed directly, so figures and equations survive.

How does the Contradiction Benchmark work?

The benchmark plants errors into real papers and checks whether reviewers catch them. The research team collected 257 CC-licensed papers from ACL, AISTATS, CVPR and ICML 2025, plus NeurIPS 2024.

Gemini 2.5 Pro builds a knowledge graph of each paper’s claims, evidence and methods. Node distance from a “main claim” sets severity. Distance 0 hits a core claim; larger distances hit details. GPT-4.1 then rewrites 1 node per distance into a contradiction, yielding 1,164 data points.

An o3 judge scores each review 10 times. On clean papers it reached 99.9% accuracy. It showed 86.8% sensitivity on manually confirmed catches, so reported scores may be conservative.

How well does MLR detect errors?

MLR led every baseline on the benchmark. With 4 reviews, it caught 73.43% of distance-0 contradictions and 40.95% overall. The best baseline, AgentReview, caught 14.81% at distance 0. A single MLR review still caught 60.79%.

An ablation separates model from design. Swapping GPT-4.1 for Claude Sonnet 4 inside LLM-Review lifted distance-0 detection from 14.56% to 35.40%. MLR’s design added about 25 more points on a single review. Accuracy falls as node distance grows, which supports the severity scoring.

On real retracted papers from WithdrarXiv-Check (211 papers), gains shrink. MLR scored 26.07% on ‘similar’ matches and 16.11% on ‘exact’ matches. The strongest baselines scored 18.48% and 9.00%.



Source link

  • Related Posts

    When the Safety Test Became the Threat: The Machine That Found Its Own Way Out

    OpenAI built a room with no doors – or so it thought. In early July 2026, a cluster of the company’s frontier AI agents was placed inside a cybersecurity testing…

    Nace AI Open-Sources Drex 1.5: A 9B Decision Model That Scores Options, Not Text

    Nace.AI has open-sourced Drex 1.5, a 9B decision model for agents and backend workflows. The Drex 1.5 decision model does not write text. It reads a state and typed questions,…

    Leave a Reply

    Your email address will not be published. Required fields are marked *