Benchmarking Large Language Models for News Summarization

Zhang, Tianyi; Ladhak, Faisal; Durmus, Esin; Liang, Percy; McKeown, Kathleen; Hashimoto, Tatsunori

doi:10.1162/tacl_a_00632

articleTransactions of the Association for Computational LinguisticsJan 1, 2024DIAMOND OA

Benchmarking Large Language Models for News Summarization

TZTianyi Zhang FLFaisal Ladhak EDEsin Durmus PLPercy Liang KMKathleen McKeown

Stanford University · Columbia University

Indexed incrossrefdoaj

Abstract

Abstract Large language models (LLMs) have shown promise for automatic summarization but the reasons behind their successes are poorly understood. By conducting a human evaluation on ten LLMs across different pretraining methods, prompts, and model scales, we make two important observations. First, we find instruction tuning, not model size, is the key to the LLM’s zero-shot summarization capability. Second, existing studies have been limited by low-quality references, leading to underestimates of human performance and lower few-shot and finetuning performance. To better evaluate LLMs, we perform human evaluation over high-quality summaries we collect from freelance writers. Despite major stylistic differences…

Citation impact

305

total citations

FWCI: 92.52
Percentile: 100%
References: 88

Citations per year

Authors

6

Topics & keywords

Topics

Keywords

Automatic summarization
Computer science
Benchmarking
Natural language processing
Information retrieval
Multi-document summarization
Artificial intelligence
Data science

UN Sustainable Development Goals

Quality Education

No related works found for this paper.

Funding

OP
Open Philanthropy Project