Acoustic Token Admixture for Joint Speaker and Content Anonymization

Abstract

Speech recordings in regulated domains cannot be shared or reused for AI development without suppressing two independent re-identification channels: biometric speaker identity and linguistically identifying content. Existing approaches address each channel in isolation and require full re-synthesis, sacrificing the in-domain acoustic characteristics that make recordings valuable. We propose a unified acoustic token-space framework that jointly anonymizes both channels via per-frame admixture of encoder-derived and phoneme-conditioned tokens, with NER-triggered span replacement preserving surrounding prosody. On the VoicePrivacy 2024 benchmark the system achieves EER 42.54%, within one point of the challenge top submission, with WER 3.73%.

Date
Oct 1, 2026 10:20 AM — 10:40 AM
Location
ICC Sydney, Australia
Avatar
Brij Mohan Lal Srivastava
Co-founder and CEO, Nijta

Co-founder and CEO of Nijta. PhD in privacy-preserving speech and co-creator of the VoicePrivacy Challenge.