Research Collection School Of Computing and Information Systems

HPS: Hard preference sampling for human preference alignment

Publication Type

Conference Proceeding Article

Version

publishedVersion

Publication Date

7-2025

Abstract

Aligning Large Language Model (LLM) responses with human preferences is vital for building safe and controllable AI systems. While preference optimization methods based on PlackettLuce (PL) and Bradley-Terry (BT) models have shown promise, they face challenges such as poor handling of harmful content, inefficient use of dispreferred responses, and, specifically for PL, high computational costs. To address these issues, we propose Hard Preference Sampling (HPS), a novel framework for robust and efficient human preference alignment. HPS introduces a training loss that prioritizes the most preferred response while rejecting all dispreferred and harmful ones. It emphasizes “hard” dispreferred responses — those closely resembling preferred ones — to enhance the model’s rejection capabilities. By leveraging a single-sample Monte Carlo sampling strategy, HPS reduces computational overhead while maintaining alignment quality. Theoretically, HPS improves sample efficiency over existing PL methods and maximizes the reward margin between preferred and dispreferred responses, ensuring clearer distinctions. Experiments on HH-RLHF and PKU-Safety datasets validate HPS’s effectiveness, achieving comparable BLEU and reward scores while greatly improving reward margins and thus reducing harmful content generation. The source code is available at https://github.com/LVLab-SMU/HPS.

Keywords

Alignment, Preference Optimization, RLHF, Large Language Models

Discipline

Programming Languages and Compilers

Research Areas

Intelligent Systems and Optimization

Areas of Excellence

Digital transformation

Publication

Proceedings of the 42nd International Conference on Machine Learning, ICML 2025, Vancouver, Canada, July 13-19

First Page

Last Page

City or Country

Vancouver, Canada

Citation

ZOU, Xiandong; LIN, Wanyu; LI, Yuchen; and ZHOU, Pan. HPS: Hard preference sampling for human preference alignment. (2025). Proceedings of the 42nd International Conference on Machine Learning, ICML 2025, Vancouver, Canada, July 13-19. 1-24.
Available at: https://ink.library.smu.edu.sg/sis_research/10401

Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-No Derivative Works 4.0 International License.

Additional URL

https://openreview.net/forum?id=hLvWwRZkok

Download

Included in

Programming Languages and Compilers Commons

COinS

Research Collection School Of Computing and Information Systems

HPS: Hard preference sampling for human preference alignment

Publication Type

Version

Publication Date

Abstract

Keywords

Discipline

Research Areas

Areas of Excellence

Publication

First Page

Last Page

City or Country

Citation

Creative Commons License

Additional URL

Included in

Search

Links

Browse

Links

Research Collection School Of Computing and Information Systems

HPS: Hard preference sampling for human preference alignment

Author

Publication Type

Version

Publication Date

Abstract

Keywords

Discipline

Research Areas

Areas of Excellence

Publication

First Page

Last Page

City or Country

Citation

Creative Commons License

Additional URL

Included in

Share

Search

Links

Browse

Links