Large language models can summarise a short story well enough to impress a professional writer — but they do so only about half the time. The other half, they fabricate character motivations, flatten subtext into surface-level observation, and copy distinctive phrasing without attribution. For any organisation relying on AI to process nuanced text — customer feedback, legal documents, qualitative research — that coin-flip reliability is a problem worth understanding.

The comprehension gap nobody measures

Most LLM benchmarks test factual recall or instruction-following. They rarely test whether models actually understand what a text means beneath its surface. A team at Columbia University designed an evaluation that does exactly that, and the results challenge comfortable assumptions about AI reading comprehension.

The researchers — Melanie Subbiah, Sean Zhang, Lydia B. Chilton, and Kathleen McKeown — recruited nine skilled creative writers (eight with writing degrees, seven previously published) and collected 25 unpublished short stories. The stories had never appeared online, meaning no model could have encountered them during training. Each writer then evaluated the AI-generated summaries of their own work, making them uniquely qualified judges.

Strategic Reality: This study used unpublished stories specifically to prevent models from drawing on memorised analysis. When LLMs encounter genuinely novel text, their comprehension drops measurably.

Three models were tested: GPT-4, Claude 2.1, and LLama-2-70B. Writers rated summaries across four dimensions — coverage, faithfulness, coherence, and analysis — using a four-point scale, and also annotated specific errors at the span level.

MetricGPT-4ClaudeLLama
Coverage3.483.172.40
Faithfulness3.122.651.92
Coherence3.523.413.08
Analysis3.403.262.76
Perfect scores (%)54%43%18%

The numbers tell a clear story. GPT-4 led on every dimension, but still received a perfect score on only 54% of summaries. Claude followed at 43%. Both models appeared fluent and competent on the surface, yet writers found substantive errors in roughly half of all outputs.

Where models actually fail

The error taxonomy is where this research becomes genuinely useful for business applications. The researchers categorised 352 span-level errors across all models and found patterns that apply well beyond fiction.

Vagueness dominates coverage failures. Across all three models, 64 of 78 coverage errors were classified as “vague” — the model covered a topic but without the specificity that makes a summary useful. In business terms, this is the difference between “the customer expressed dissatisfaction” and “the customer threatened to cancel their subscription over billing errors.”

Faithfulness errors cluster around feelings and actions. Models consistently misread characters’ emotional states and misrepresented what characters actually did. Claude, for instance, interpreted physical descriptions like “He looked nice” and “I felt warm in the cheeks” as signs of romantic attraction between the wrong characters. In another case, a model stated a narrator “decided to stay at the summit” when the text actually implied the narrator died there.

Critical Context: Faithfulness errors are not random hallucinations. They follow predictable patterns — models impose conventional interpretive frameworks on unconventional material. A character described as having a “difficult relationship” with their mother was actually close to their mother, whose near-death from cancer made circumstances difficult.

Unsupported analysis is rampant. Of 104 analysis errors, models drew conclusions the stories simply did not support. One writer commented that the summary’s analysis was “either incorrectly interpreted or missing” when it came to subtext. Models produced analysis that sounded thoughtful but rested on misreadings of the source material.

Attribution is absent. GPT-4 copied an average of 5.72 words of exact-match phrasing from the original stories without quotation marks. All three models showed significant n-gram overlap with source texts — 22.53% of bigrams for GPT-4, 19.80% for Claude. Writers flagged this as bordering on plagiarism.

What makes text hard for AI to understand

The researchers used Genette’s model of narrative elements to identify which story characteristics predict poor AI performance. Three factors stood out.

Unreliable narrators break models consistently. Stories with narrators who say one thing and mean another produced lower scores across every model. One narrator described himself as “practical” and “logical” while his behaviour showed the opposite. GPT-4 took the narrator at his word, summarising the character as someone who “shares a desire for a simple life” based solely on the narrator’s self-description — ignoring textual evidence to the contrary.

Strategic Insight: Unreliable narrators are not just a literary device. Customer communications, employee surveys, and negotiation transcripts all contain statements where the speaker’s intended meaning differs from the literal words. Models that cannot detect this gap will produce misleading analysis.

Non-linear timelines confuse open-source models. Stories with flashbacks and disrupted chronology produced significantly lower scores for LLama (average 2.50 versus 3.05 for linear stories). GPT-4 and Claude handled non-linear structure somewhat better, but the gap indicates that document structure affects comprehension quality.

Story length is not the problem people expect. Long-context models (GPT-4 and Claude) summarised stories up to 10,000 tokens with no significant quality drop compared to shorter stories. The challenge is not processing capacity — it is interpretive depth.

The LLM-as-judge problem

Perhaps the most strategically important finding: LLM self-evaluation does not work for this kind of task.

When GPT-4 and Claude were asked to rate the same summaries that writers had evaluated, their scores showed little correlation with human judgements. Claude rated 85% of summaries as 4 out of 4 on faithfulness, compared to the writers’ average of 27%. GPT-4 overestimated its own faithfulness scores (64% perfect versus the writers’ 44%).

Warning: ⚠️ Organisations using LLM-as-judge pipelines for quality assurance on nuanced content should treat this finding seriously. Models rate their own outputs — and each other’s — as far more accurate than human experts do.

Both models consistently overestimated coherence and underestimated coverage. The gap between model confidence and human assessment was largest precisely where it matters most: faithfulness and analytical depth.

Evaluation sourceFaithfulness (% rated 4/4)Analysis (% rated 4/4)
Writers (ground truth)27% average39% average
Claude self-rating85%85%
GPT-4 self-rating64%

Standard automatic metrics performed no better. ROUGE, BERTScore, and specialised faithfulness metrics (AlignScore, UniEval, MiniCheck) showed weak or no correlation with writer ratings. The researchers found no existing metric that reliably predicts summary quality for narrative text.

Who should pay attention and why

This research matters beyond the NLP community because it quantifies a failure mode that many business applications assume away.

Content analysis teams using LLMs to process qualitative data — interview transcripts, open-ended survey responses, user research — face the same pattern of errors. Models will produce plausible-sounding analysis that misses the subtext human readers would catch.

Legal and compliance teams relying on AI for document review should note the faithfulness error patterns. Models misrepresent causal relationships (why something happened) and character states (what someone intended) at rates that would be unacceptable in legal analysis.

Implementation Note: The 50% excellent-summary rate applies to 2023-era models (GPT-4, Claude 2.1). Newer models may perform differently, but the underlying failure patterns — subtext blindness, normative bias, attribution absence — reflect architectural limitations rather than scale problems.

Publishing and media organisations should note the plagiarism concern. Models copy distinctive phrasing without attribution, and the copying is substantial enough that writers flagged it unprompted.

AI evaluation teams should reconsider LLM-as-judge approaches for any task requiring nuanced comprehension. The study demonstrates that model self-assessment is systematically miscalibrated on exactly the dimensions that matter for content quality.

Building safeguards around known weaknesses

The research points toward specific mitigations rather than general caution.

Priority actions for organisations processing nuanced text:

  1. Audit for subtext sensitivity. Test your LLM pipeline on inputs where the literal meaning differs from the intended meaning. Customer complaints, diplomatic communications, and qualitative research frequently contain this kind of gap.

  2. Do not rely on LLM-as-judge for quality assurance. The correlation between model self-ratings and expert human ratings is too weak to serve as a reliable filter. Human review remains necessary for high-stakes content processing.

  3. Watch for normative bias in analysis. Models impose conventional interpretive frameworks. One summary assumed a heteronormative reading of a relationship that the text did not support. Another imposed a “redemption arc” narrative where the writer intended ambiguity.

  4. Check for unattributed copying. If your application summarises proprietary or copyrighted text, audit outputs for n-gram overlap. The models in this study copied significant amounts of source phrasing without quotation marks.

Take Action: Before deploying LLMs for content analysis, create a test set of 20-30 texts containing subtext, irony, or ambiguity. Have domain experts evaluate the model outputs. If accuracy falls below your threshold, implement human-in-the-loop review for those categories.

For teams building evaluation pipelines:

The study suggests that coverage is the one dimension where automatic metrics show moderate correlation with human judgement. Faithfulness, coherence, and analysis remain essentially unmeasurable by automated means for nuanced text. Design your evaluation infrastructure accordingly — automate what can be automated, but budget for human review where it cannot.

The real cost of surface-level comprehension

The writers in this study provided a perspective that benchmarks miss. On the best summaries, they wrote: “it did analysis that even I — the writer — hadn’t done! Very clever.” On the worst: “completely misses the fundamental thread” and “makes use of stock phrases or stand-ins rather than accurate, specific summary.”

Reality Check: Models that appear to understand text — producing fluent, well-structured summaries — may be performing sophisticated pattern-matching rather than genuine comprehension. The difference only becomes visible when domain experts evaluate the output.

That gap between appearance and reality is the core business risk. A summary that sounds competent but misrepresents emotional dynamics, causal relationships, or authorial intent is worse than no summary at all — because it creates false confidence in flawed analysis.

The Columbia team’s methodology offers a template for any organisation that wants to measure this risk: work with domain experts, test on material the model has never seen, and evaluate not just whether the output is fluent but whether it is right.


Source: Subbiah, M., Zhang, S., Chilton, L. B., & McKeown, K. (2024). Reading Subtext: Evaluating Large Language Models on Short Story Summarization with Writers. Columbia University. Published at TACL 2024.

Strategic analysis by Resultsense — Making sense of AI in the UK.