Time-Frequency Consistency Learning for Robust Speech Deepfake Detection
5 hours ago
- Current speech deepfake detection (SDD) robustness evaluation is limited to additive noise, not addressing complex distortions from acoustic front-end (AFE) processing pipelines.
- A unified AFE pipeline (including echo cancellation, noise suppression, automatic gain control, and voice activity detection) significantly degrades detection performance due to nonlinear and time-frequency coupled distortions.
- The proposed Time-Frequency Consistency Learning (TFCL) framework uses an attention-driven soft alignment mechanism and frequency-domain structural consistency constraints to learn invariant spoofing representations stable before and after AFE processing.
- Extensive experiments show that TFCL effectively mitigates performance degradation caused by AFE, improving SDD robustness in real-world scenarios.