Перейти к содержимому

[SPSC Webinar 20260820] SALMONN-Guard, SH-Bench, and HoliAntiSpoof!

SPSC Webinar

0:00 / 0:00

[SPSC Webinar 20260820] SALMONN-Guard, SH-Bench, and HoliAntiSpoof!

8 просмотров · 2 недели назад
SPSC Webinar
11 подписчиков
8 просмотров · 2 недели назад
Learn more about SALMONN-Guard, SH-Bench, and HoliAntiSpoof! 📌 Title: When Hearing Everything Becomes a Risk: Safety and Privacy in Audio LLMs 🗓 Time: Wed., Aug. 20, 15:00 - 15:30 CET 🎙 Speaker: Dr. Guangzhi Sun, Trinity College 📌 Title: Unifying Speech Editing Detection and Content Localization via Prior-Enhanced Audio LLMs 🗓 Time: Wed., Aug. 20, 15:30 - 16:00 CET 🎙 Speaker: Dr. Xuenan Xu, Shanghai AI Lab Talk 1: Abstract: Complex audio introduces new safety and privacy challenges for audio LLMs. This talk examines safety vulnerabilities arising from speech–audio composition and privacy risks involving bystanders in a multi-speaker audio stream. From a safety perspective, harmful content can be concealed within overlapping speech, non-speech audio, or multi-speaker dialogue, exposing limitations of existing text-only safeguards. SALMONN-Guard addresses these vulnerabilities through audio-aware safety assessment. From a privacy perspective, real-world audio may contain speech from bystanders who do not intend to interact with the system. Selective hearing provides a framework for protecting such bystander information while preserving the model’s ability to understand the intended speaker. Together, these studies highlight the need for safety and privacy mechanisms designed specifically for complex, multi-party audio environments. Speaker: Guangzhi Sun (Brian) is a Research Fellow at Trinity College, Cambridge, specialising in the capability and robustness of multi-modal LLMs, especially audio-visual understanding. With a PhD from the University of Cambridge, Brian bridges the gap between academic rigor and industrial scale, having previously contributed to teams at Google Brain, ByteDance, and PolyAI. His recent research--published at ICML, EMNLP, and ACL--tackles critical challenges in AI safety, such as bystander privacy and multi-modal jailbreaking. Talk 2: Abstract: Recent advances in speech synthesis and editing have made speech spoofing increasingly challenging, yet most existing anti-spoofing methods focus on binary real/fake classification. In this talk, I will introduce HoliAntiSpoof, an audio large language model (ALLM) framework for holistic speech anti-spoofing analysis. Rather than treating spoofing detection as a single classification task, HoliAntiSpoof jointly analyzes the authenticity of speech, spoofing methods, manipulated regions, and the semantic influence of spoofed content within a unified text generation framework. To fill the gap of semantic influence analysis of spoofed words, a new dataset, DailyTalkEdit, is curated for realistic conversational speech manipulations. Trained on large-scale anti-spoofing datasets with various spoofing operations, HoliAntiSpoof outperforms conventional anti-spoofing methods across in-domain and out-of-domain settings. In-context learning further enables the model to adapt to unseen spoofing scenarios and languages without additional model training. These studies highlights the potential of audio LLMs for more comprehensive, interpretable, and generalizable speech security analysis. Speaker: Xuenan Xu received the B.S. and Ph.D. degrees from Shanghai Jiao Tong University. His research focuses on audio and speech processing, with particular interests in audio large language models and general audio generation. His work explores building general-purpose models that can understand and generate diverse forms of audio, including speech, music, and environmental sounds. His work has been published in major conferences and journals in speech and audio processing.