Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization

Jul 02, 2025

Keyan Jin, Yapeng Wang, Leonel Santos, Tao Fang, Xu Yang, Sio Kei Im, Hugo Gonçalo Oliveira

Figure 1 for Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization

Figure 2 for Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization

Figure 3 for Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization

Figure 4 for Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization

Share this with someone who'll enjoy it:

Abstract:Dialogue summarization is a challenging task with significant practical value in customer service, meeting analysis, and conversational AI. Although large language models (LLMs) have achieved substantial progress in summarization tasks, the performance of step-by-step reasoning architectures-specifically Long Chain-of-Thought (CoT) implementations such as OpenAI-o1 and DeepSeek-R1-remains unexplored for dialogue scenarios requiring concurrent abstraction and conciseness. In this work, we present the first comprehensive and systematic evaluation of state-of-the-art reasoning LLMs and non-reasoning LLMs across three major paradigms-generic, role-oriented, and query-oriented dialogue summarization. Our study spans diverse languages, domains, and summary lengths, leveraging strong benchmarks (SAMSum, DialogSum, CSDS, and QMSum) and advanced evaluation protocols that include both LLM-based automatic metrics and human-inspired criteria. Contrary to trends in other reasoning-intensive tasks, our findings show that explicit stepwise reasoning does not consistently improve dialogue summarization quality. Instead, reasoning LLMs are often prone to verbosity, factual inconsistencies, and less concise summaries compared to their non-reasoning counterparts. Through scenario-specific analyses and detailed case studies, we further identify when and why explicit reasoning may fail to benefit-or even hinder-summarization in complex dialogue contexts. Our work provides new insights into the limitations of current reasoning LLMs and highlights the need for targeted modeling and evaluation strategies for real-world dialogue summarization.

View paper on

Share this with someone who'll enjoy it:

Title:Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization

Paper and Code