Toxic content propagation across networked communication
platforms constitutes an emerging cybersecurity
challenge, as real-time audio and image channels introduce
attack surfaces that bypass conventional text-only defences. Automated
content moderation in online communication platforms
increasingly requires coverage beyond text, as toxic content is
frequently delivered through audio and image channels that textonly
systems cannot address. This paper presents a dual-modality
toxic content detection and moderation pipeline operating across
audio and image inputs in real time. For audio, the Roblox
voice-safety-classifier-v2, a WavLM transformer pretrained on
over 100,000 hours of real gaming voice chat, generates sixdimensional
toxicity probability scores mapped to the Jigsaw
taxonomy via a novel cross-taxonomy semantic alignment, with
faster-whisper providing word-level timestamps for surgical muting
of precisely the toxic speech segments rather than entire
clips. For image moderation, CLIP ViT-B/32 encodes images
against toxicity-describing natural language prompts to produce
a nine-dimensional feature vector, with flagged content reposted
as blurred spoilers. Proximal Policy Optimisation reinforcement
learning agents trained with an asymmetric reward structure
penalising false negatives over false positives achieve 97.5%
accuracy on unseen audio evaluation data with 91.0% reward efficiency,
and 87.1% reward efficiency on image data, significantly
outperforming rule-based, random, always-mute, and alwaysallow
baseline policies. A per-user cross-modal trust score system
with progressive escalation is deployed as a real-time automated
Discord moderation bot validated through live user interactions.
07月30日
2026
08月01日
2026
注册截止日期
初稿截稿日期
2026年07月30日 印度 Trichy
2026 International Conference on Networks Computers and Communications
发表评论