The gains do not make up for the losses: A comprehensive evaluation for safety alignment of large language models via machine unlearning
Publication Type
Journal Article
Version
publishedVersion
Publication Date
2-2026
Abstract
Machine Unlearning (MU) has emerged as a promising technique for aligning large language models (LLMs) with safety requirements to steer them forgetting specific harmful contents. Despite the significant progress in previous studies, we argue that the current evaluation criteria, which solely focus on safety evaluation, are actually impractical and biased, leading to concerns about the true effectiveness of MU techniques. To address this, we propose to comprehensively evaluate LLMs after MU from three aspects: safety, over-safety, and general utility. Specifically, a novel benchmark MuBench with 18 related datasets is first constructed, where the safety is measured with both vanilla harmful inputs and 10 types of jailbreak attacks. Furthermore, we examine whether MU introduces side effects, focusing on over-safety and utility-loss. Extensive experiments are performed on 3 popular LLMs with 7 recent MU methods. The results highlight a challenging trilemma in safety alignment without side effects, indicating that there is still considerable room for further exploration. MuBench serves as a comprehensive benchmark, fostering future research on MU for safety alignment of LLMs.
Keywords
Large language models, Machine unlearning, Safety alignment
Discipline
Artificial Intelligence and Robotics | Information Security
Research Areas
Intelligent Systems and Optimization
Areas of Excellence
Digital transformation
Publication
Frontiers of Computer Science
Volume
20
Issue
2
First Page
1
Last Page
25
ISSN
2095-2228
Identifier
10.1007/s11704-024-41099-x
Publisher
Springer
Citation
ZHAO, Weixiang; HU, Yulin; SUI, Xingyu; LI, Zhuojun; DENG, Yang; ZHAO, Yanyan; QIN, Bing; and CHE, Wanxiang.
The gains do not make up for the losses: A comprehensive evaluation for safety alignment of large language models via machine unlearning. (2026). Frontiers of Computer Science. 20, (2), 1-25.
Available at: https://ink.library.smu.edu.sg/sis_research/11297
Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-No Derivative Works 4.0 International License.
Additional URL
https://doi.org/10.1007/s11704-024-41099-x