首页 > AI前沿 > LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization

LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization

arXiv自然语言 2026-06-04 12:12 10 阅读 查看原文

Although context windows have expanded significantly in recent years, hallucinations in long-context summarization remain a challenge.

Long novels are better suited than news or papers for researching these hallucinations, due to their intrinsic information and detailed descriptions of events and dialogues.

However, current research lacks a multi-scale benchmark for hallucination detection in long-context novel summarization and does not fully explore how hallucinations change as the context grows longer.

In this Study

In this study, we propose LongNovel, a multi-scale long-context bilingual (Chinese and English) novel benchmark for hallucination detection.

This benchmark is constructed from 29 Chinese novels (ranging from 16k to 100k tokens) and chapter-level data from the BookSum dataset.

We design 8 hallucination types and employ a combination of Multi-Model Arbitration and Entity-Referenced Hallucination Generation to ensure both data authenticity and a balanced distribution of hallucination categories.

Furthermore, we manually revise the content in the test set to guarantee data reliability.

Experimental Results

Extensive experimental results demonstrate that LongNovel is a challenging benchmark.

We release LongNovel for future research.

https://github.com/BDML-lab/LongNovel