Brief
SAGE scores literary quality in six layers, and finds the interpretive gap is roughly double the emotional one
A paper describes SAGE, a six-layer framework that splits rule-based checks of observable text properties from LLM-judged interpretive qualities. The evidence covers the paper's reported measurements only: no rubric text, code or release is described, so treat the framework as a method to inspect rather than a tool to run.

The claim is that existing natural-language generation metrics cannot measure interpretive dimensions — cultural representation, emotional depth, philosophical engagement — and SAGE is built to fill that gap. It keeps rule-based assessment of observable textual properties separate from LLM-based evaluation of interpretive qualities drawn from cultural theory, affect theory and existentialist philosophy.
Each interpretive layer is judged through multi-round iterative LLM evaluation with independent cross-validation, reporting 98.8% convergence and more than 94% inter-rater agreement, stable across evaluator models. The validation covers 600 evaluations across 100 short stories. The central finding is a boundary: emotional-psychological representation approaches human levels, while cultural critique and philosophical depth show roughly double the gap. LLM-generated narratives scored below even commercial genre fiction on all three layers.
Our reading
This matters because it is a measurement claim about what narrative generation can and cannot do, not another headline score: the split between pattern-reproducible capacities and stance-requiring ones gives evaluators a place to put the failure rather than averaging it away. Anyone using generated text for cultural or argumentative work — and anyone building an evaluation harness that currently…
What to do or watch
A bounded next step is to copy the structural move rather than the scores: separate properties you can check mechanically from judgments you cannot, and cross-validate the interpretive judgments against more than one evaluator model before trusting a number. The unresolved question is whether the layer definitions and rubric are published anywhere usable, since this evidence does not say.
Source details and supporting facts
Each line is stated by the page named above it.
Stated by arXiv
- Each interpretive layer is assessed through multi-round iterative LLM evaluation with independent cross-validation, achieving 98.8% convergence and more than 94% inter-rater agreement, stable across evaluator models.
- The framework was validated on 600 evaluations across 100 short stories.
- Emotional-psychological representation approaches human levels, while cultural critique and philosophical depth exhibit approximately double the gap.
- LLM-generated narratives score below even commercial genre fiction on all three layers.
Sources
- arXivText stored 16 September 2026
How this story was checked. Written from the 1 page listed above, stored 16 September 2026; claims checked against that stored text on 16 September 2026.
What that means
- 4 of 5 reported statements were confirmed against the page that carries them; the rest were removed rather than published.
- Figures in the text were required to appear in the stored source text: yes. Identifiers: yes.
- The check reads stored text only: no claim rests on a fresh look that did not happen.
- Where the reporting was silent, the text says so instead of filling the gap.