Publication Type
Conference Proceeding Article
Version
acceptedVersion
Publication Date
6-2026
Abstract
Large language models (LLMs) with reasoning capabilities have fueled a compelling narrative that reasoning universally improves performance across language tasks. We test this claim through a comprehensive evaluation of 504 configurations across seven model families—including adaptive, conditional, and reinforcement learning-based reasoning architectures—on sentiment analysis datasets of varying granularity (binary, five-class, and 27-class emotion). Our findings reveal that reasoning effectiveness is strongly task-dependent, challenging prevailing assumptions: (1) Reasoning shows task-complexity dependence—binary classification degrades up to -19.9 F1% points (pp), while 27-class emotion recognition gains up to +16.0 pp; (2) Distilled reasoning variants underperform base models by 3–18 pp on simpler tasks, though few-shot prompting enables partial recovery; (3) Few-shot learning improves over zero-shot in most cases regardless of model type, with gains varying by architecture and task complexity; (4) Pareto frontier analysis shows base models dominate efficiency-performance trade-offs, with reasoning justified only for complex emotion recognition despite 2.1X–54X computational overhead. We complement these quantitative findings with qualitative error analysis revealing that reasoning degrades simpler tasks through systematic over-deliberation, offering mechanistic insight beyond the high-level overthinking hypothesis.
Discipline
Artificial Intelligence and Robotics | Databases and Information Systems
Research Areas
Intelligent Systems and Optimization
Areas of Excellence
Digital transformation
Publication
Proceedings of the 30th Pacific-Asia Conference on Knowledge Discovery and Data Mining, PAKDD 2026, Hong Kong, China, June 9-12
First Page
288
Last Page
299
Identifier
10.1007/978-981-92-1468-6_18
Publisher
Springer
City or Country
Cham
Citation
HUANG, Donghao and WANG, Zhaoxia.
Task complexity matters: An empirical study of reasoning in LLMs for sentiment analysis. (2026). Proceedings of the 30th Pacific-Asia Conference on Knowledge Discovery and Data Mining, PAKDD 2026, Hong Kong, China, June 9-12. 288-299.
Available at: https://ink.library.smu.edu.sg/sis_research/11232
Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-No Derivative Works 4.0 International License.
Additional URL
https://doi.org/10.1007/978-981-92-1468-6_18