article
As neural vocoders increasingly produce perceptually seamless deepfakes, conventional detectors relying on fragile, localized artifacts fail against real-world degradations such as lossy compression. To address this, we propose the Hybrid Graph Siamese Network (GSNet), which models the global structural integrity of audio signals rather than relying on local anomalies. Operating within a Siamese metric learning framework, GSNet integrates a CNN backbone for feature extraction with a Graph Neural Network (GNN) refiner to capture long-range, non-Euclidean acoustic dependencies. Empirical validation across multiple benchmarks demonstrates superior robustness: GSNet achieves an Equal Error Rate (EER) of 0.929% on the ASVspoof 2021 (DF) task and 0.536% on the ASVspoof 5 (Track 1) closed condition. Furthermore, the model demonstrates strong cross-domain generalization, attaining 0.048% EER on the WaveFake datasets. These results confirm that prioritizing structural consistency over artifact detection significantly enhances resilience against compression and unseen generative attacks.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.36227/techrxiv.177130683.36376095/v1
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.