Brief
New reporting framework separates implementation from evaluation in AI-assisted systematic reviews
A new framework called PRISMA-LLM has been proposed to improve reporting for AI-assisted systematic reviews. It separates implementation disclosure from consequence-sensitive evaluation and limitation reporting, based on an analysis of 888 review-automation papers.
The framework, PRISMA-LLM, is described as empirically grounded. Its authors analyzed SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations. They found that since 2023, 38.0% of software/product papers reported no evaluation, compared with 9.3% of LLM papers. Reporting coverage increased with LLM workflow complexity, yet 52% of positive-only LLM evaluations still reported an unmet reliability or performance requirement. The framework aims to separate implementation disclosure from consequence-sensitive evaluation and limitation reporting.
Source details and supporting facts
Each line is stated by the page named above it.
Stated by arXiv
- PRISMA-LLM is an empirically grounded framework separating implementation disclosure from consequence-sensitive evaluation and limitation reporting.
- The analysis used SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations.
- Since 2023, 38.0% of software/product papers reported no evaluation, compared with 9.3% of LLM papers.
- 52% of positive-only LLM evaluations still reported an unmet reliability or performance requirement.
Sources
- arXivText stored 16 September 2026
How this story was checked. Written from the 1 page listed above, stored 16 September 2026; claims checked against that stored text on 16 September 2026.
What that means
- 4 of 4 reported statements were confirmed against the page that carries them; the rest were removed rather than published.
- Figures in the text were required to appear in the stored source text: yes. Identifiers: yes.
- The check reads stored text only: no claim rests on a fresh look that did not happen.
- Where the reporting was silent, the text says so instead of filling the gap.