Publication Type
Conference Proceeding Article
Version
publishedVersion
Publication Date
7-2026
Abstract
The widespread availability of large-scale code datasets has accelerated the development of code large language models (CodeLLMs), raising concerns about unauthorized dataset usage. Dataset poisoning offers a proactive defense by reducing the utility of such unauthorized training. However, existing poisoning methods often require full-dataset poisoning and introduce transformations that break code compilability. In this paper, we introduce FunPoison, a functionality-preserving poisoning approach that injects short, compilable weak-use fragments into executed code paths. FunPoison leverages reusable statement-level templates with automatic repair and conservative safety checking to ensure side-effect freedom, while a type-aware synthesis module preserves type correctness, suppresses static-analysis warnings, and improves stealth. Extensive experiments across multiple CodeLLMs and code-generation benchmarks show that FunPoison achieves effective poisoning by contaminating only 10% of the dataset, while maintaining 100% compilability and functional correctness. FunPoison also remains robust against advanced code sanitization techniques, including detection, purification, rewriting, static-analysis, and formatting defenses.
Discipline
Artificial Intelligence and Robotics | Software Engineering
Research Areas
Software and Cyber-Physical Systems
Areas of Excellence
Digital transformation
Publication
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), San Diego, California, July 2-7
First Page
31284
Last Page
31303
Identifier
10.18653/v1/2026.findings-acl.1564
Publisher
ACL
City or Country
San Diego, California
Citation
XIAO, Yuan; CHEN, Yuchen; WANG, Jiaming; SONG, Wei; SUN, Jun; MA, Shiqing; MU, Yanzhou; ZHAI, Juan; FANG, Chunrong; DONG, Jin Song; and CHEN, Zhenyu.
Train in vain: Functionality-preserving poisoning to prevent unauthorized use of code datasets. (2026). Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), San Diego, California, July 2-7. 31284-31303.
Available at: https://ink.library.smu.edu.sg/sis_research/11197
Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-No Derivative Works 4.0 International License.
Additional URL
https://doi.org/10.18653/v1/2026.findings-acl.1564