BriefPulse Practical AI · Working notes on AI you can actually use. RSS · BriefPulse network
BriefPulse Practical AI

What changed in AI, what it is useful for, and what you can do with it.

16 September 2026

Brief

New reporting framework separates implementation from evaluation in AI-assisted systematic reviews

A new framework called PRISMA-LLM has been proposed to improve reporting for AI-assisted systematic reviews. It separates implementation disclosure from consequence-sensitive evaluation and limitation reporting, based on an analysis of 888 review-automation papers.

New reporting framework separates implementation from evaluation in AI-assisted systematic reviews:
Original graphic. Every figure in it is stated in the reporting; the sources are listed below this article.

The framework, PRISMA-LLM, is described as empirically grounded. Its authors analyzed SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations. They found that since 2023, 38.0% of software/product papers reported no evaluation, compared with 9.3% of LLM papers. Reporting coverage increased with LLM workflow complexity, yet 52% of positive-only LLM evaluations still reported an unmet reliability or performance requirement. The framework aims to separate implementation disclosure from consequence-sensitive evaluation and limitation reporting.

Source details and supporting facts

Each line is stated by the page named above it.

Stated by arXiv

  • PRISMA-LLM is an empirically grounded framework separating implementation disclosure from consequence-sensitive evaluation and limitation reporting.
  • The analysis used SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations.
  • Since 2023, 38.0% of software/product papers reported no evaluation, compared with 9.3% of LLM papers.
  • 52% of positive-only LLM evaluations still reported an unmet reliability or performance requirement.

Sources

  1. arXivText stored 16 September 2026

How this story was checked. Written from the 1 page listed above, stored 16 September 2026; claims checked against that stored text on 16 September 2026.

What that means
  • 4 of 4 reported statements were confirmed against the page that carries them; the rest were removed rather than published.
  • Figures in the text were required to appear in the stored source text: yes. Identifiers: yes.
  • The check reads stored text only: no claim rests on a fresh look that did not happen.
  • Where the reporting was silent, the text says so instead of filling the gap.

More from Practical AI