Semantic–emotional matching-based detection for partially fake audio

Publication Type

Conference Proceeding Article

Publication Date

7-2026

Abstract

Existing partially fake audio detection methods rely heavily on semantic cues, leading to limited sensitivity to emotional inconsistencies and poor frame-level discrimination. To address this issue, this paper proposes a semantic–emotional matching-based dual-stream detection framework for frame-level partially fake audio detection. The proposed framework learns context-aware representations from parallel semantic and emotional streams and models their intrinsic correspondence through structured matching within a unified representation space. A semantic–emotional matching-driven local enhancement mechanism is further introduced to learn matching patterns between the two streams. Based on their temporal similarity, the model adaptively enhances local features, preserving coherent patterns in genuine speech while highlighting inconsistencies in fake frames. By jointly characterizing semantic irregularities and emotional deviations, the proposed framework produces more discriminative frame-level representations. Experimental results on the PartialSpoof and ADD 2023 datasets demonstrate improved performance over representative baselines in frame-level partially fake audio detection.

Keywords

Local Attention, Partially Fake Audio Detection, Semantic-Emotional Matching

Discipline

Artificial Intelligence and Robotics | Information Security

Research Areas

Intelligent Systems and Optimization

Publication

Proceedings of the 22nd International Conference on Intelligent Computing, ICIC 2026, Toronto, Canada, July 22-26

First Page

531

Last Page

542

ISBN

9789819234226

Identifier

10.1007/978-981-92-3423-3_43

Publisher

Springer

City or Country

Cham

Additional URL

https://doi.org/10.1007/978-981-92-3423-3_43

This document is currently not available here.

Share

COinS