CoolFace
Datasetpublic

lab260/RuASD

RuASD: Russian Anti-Spoofing Dataset RuASD is a public Russian-language speech anti-spoofing dataset designed for developing and benchmarking audio deepfake detection systems. It combines spoofed utterances generated by 37 Russian-capable speech synthesis systems with bona fide recordings curated from multiple heterogeneous Russian speech corpora. In addition to clean audio, the dataset supports robustness-oriented evaluation through reproducible perturbations such as reverberation… See the full description on the dataset page: https://huggingface.co/datasets/lab260/RuASD.

sourceHugging Facecc-by-nc-sa-4.0updated 3mo agoView on Hugging Face
1likes340downloads
Dataset Card

RuASD: Russian Anti-Spoofing Dataset

RuASD is a public Russian-language speech anti-spoofing dataset designed for developing and benchmarking audio deepfake detection systems. It combines spoofed utterances generated by 37 Russian-capable speech synthesis systems with bona fide recordings curated from multiple heterogeneous Russian speech corpora. In addition to clean audio, the dataset supports robustness-oriented evaluation through reproducible perturbations such as reverberation, additive noise, and codec-based channel degradation.

Models: ESpeech, F5-TTS, VITS, Piper, TeraTTS, MMS TTS, VITS2, GPT-SoVITS, CoquiTTS, XTSS, Fastpitch, RussianFastSpeech, Bark, GradTTS, FishTTS, Pyttsx3, RHVoice, Silero, Fairseq Transformer, SpeechT5, Vosk-TTS, EdgeTTS, VK Cloud, SaluteSpeech, ElevenLabs

Overview

  • —Purpose: Benchmark and develop Russian-language anti-spoofing and audio deepfake detection systems, with a focus on robustness to realistic channel and post-processing distortions.
  • —Content: Bona fide speech from multiple open Russian speech corpora and synthetic speech generated by 37 Russian-capable TTS and voice-cloning systems.
  • —Structure:
  • —Audio: .wav files
  • —Metadata: JSON with the fields sample_id, label, group, subset, augmentation, filename, audio_relpath, source_audio, metadata_source, source_type, mos_pred, noi_pred, dis_pred, col_pred, loud_pred, cer, duration, speakers, model, transcribe, true_lines, transcription, ground_truth, and ops.
FieldDescription
sample_idSample ID
labelreal or fake
groupSample group - raw or augmented
subsetsource subset name, e.g. OpenSTT, GOLOS, or ElevenLabs
augmentationApplied augmentation
filenameAudio filename
audio_relpathRelative path to audio
source_audioOriginal audio for augmented sample
metadata_sourceMetadata source
source_typeSource type - tts, real_speech or augmented_audio
mos_predPredicted MOS
noi_predPredicted noisiness
dis_predPredicted discontinuity
col_predPredicted coloration
loud_predPredicted loudness
cerCharacter error rate
durationDuration in seconds
speakersSpeaker info
modelspecific checkpoint or voice used for generation, e.g. ESpeech-TTS-1_RL-V1, xtts-ru-ipa, or ru-RU-DmitryNeural
transcribeAutomatic transcription
true_linesSource text
transcriptionAutomatic transcription
ground_truthReference text
opsProcessing operations

Statistics

  • —Number of TTS systems: 37
  • —Total spoof hours: 691.68
  • —Total bona-fide hours: 234.07

Table 4. Antispoofing models on clean data

ModelAccPrRecF1RAUCEERt-DCF
AASIST30.769±0.00060.683±0.0010.769±0.00060.724±0.0010.841±0.00060.231±0.00060.702±0.002
Arena-1B0.812±0.0010.736±0.0010.812±0.0010.772±0.0010.887±0.00050.188±0.001<u>0.385±0.001</u>
Arena-500M0.801±0.0010.722±0.0010.801±0.0010.760±0.0010.864±0.00050.199±0.0010.655±0.002
Nes2Net0.689±0.00070.589±0.0010.689±0.00070.634±0.00080.779±0.00070.311±0.00070.696±0.001
Res2TCNGaurd0.627±0.0010.520±0.0010.627±0.0010.569±0.0010.691±0.0010.373±0.0010.918±0.001
ResCapsGuard0.677±0.0010.575±0.0010.677±0.0010.622±0.0010.718±0.0010.323±0.0010.896±0.001
SLS with XLS-R0.779±0.0010.700±0.0010.779±0.0010.737±0.0010.859±0.0010.221±0.0010.650±0.001
Wav2Vec 2.00.772±0.00060.687±0.0010.772±0.00060.727±0.0010.850±0.00060.228±0.00060.558±0.002
TCM-ADD<u>0.857±0.001</u><u>0.797±0.001</u><u>0.859±0.001</u><u>0.827±0.001</u><u>0.914±0.0004</u><u>0.143±0.001</u>0.424±0.001
Spectra-00.9620.9420.9620.9520.9850.0380.124

Download

Using Datasets

python
from datasets import load_dataset

ds = load_dataset("MTUCI/RuASD")
print(ds)

Using Datasets with streaming mode

python
from datasets import load_dataset

ds = load_dataset("MTUCI/RuASD", streaming=True)
small_ds = ds.take(1000)

print(small_ds)

Contact

Citation

@unpublished{ruasd2026,
  author = {},
  title = {},
  year = {}
}

TTS and VC models

ModelLink
Espeech Podcasterhttps://hf.co/ESpeech/ESpeech-TTS-1_podcaster
Espeech RL-V1https://hf.co/ESpeech/ESpeech-TTS-1_RL-V1
Espeech RL-V2https://hf.co/ESpeech/ESpeech-TTS-1_RL-V1
Espeech SFT-95khttps://hf.co/ESpeech/ESpeech-TTS-1_SFT-95K
Espeech SFT-256khttps://hf.co/ESpeech/ESpeech-TTS-1_SFT-256K
F5-TTS checkpointhttps://hf.co/Misha24-10/F5-TTS_RUSSIAN
F5-TTS checkpointhttps://hf.co/hotstone228/F5-TTS-Russian
VITS checkpointhttps://hf.co/joefox/ttsvitsru_hf
PiperTTShttps://github.com/rhasspy/piper
TeraTTS-natashahttps://hf.co/TeraTTS/natasha-g2p-vits
TeraTTS-girl_nicehttps://hf.co/TeraTTS/girl_nice-g2p-vits
TeraTTS-gladoshttps://hf.co/TeraTTS/glados-g2p-vits
TeraTTS-glados2https://hf.co/TeraTTS/glados2-g2p-vits
MMShttps://hf.co/facebook/mms-tts-rus
VITS checkpointhttps://hf.co/utrobinmv/ttsrufreehfvitslowmultispeaker
VITS checkpointhttps://hf.co/utrobinmv/ttsrufreehfvitshighmultispeaker
VITS2 checkpointhttps://hf.co/frappuccino/vits2runatasha
GPT-SoVITS checkpointhttps://hf.co/alphacep/vosk-tts-ru-gpt-sovits
CoquiTTShttps://hf.co/coqui/XTTS-v2
XTTS checkpointhttps://hf.co/NeuroDonu/RU-XTTS-DonuModel
XTTS checkpointhttps://hf.co/omogr/xtts-ru-ipa
Fastpitch IPAhttps://hf.co/bene-ges/ttsruipafastpitchruslan
Fastpitch BERT g2phttps://hf.co/bene-ges/rug2pipabertlarge
RussianFastPitchhttps://github.com/safonovanastya/RussianFastPitch
Barkhttps://hf.co/suno/bark-small
GradTTShttps://github.com/huawei-noah/Speech-Backbones/tree/main/Grad-TTS
FishTTShttps://hf.co/fishaudio/fish-speech-1.5
Pyttsx3https://github.com/nateshmbhat/pyttsx3
RHVoicehttps://github.com/RHVoice/RHVoice
Silerohttps://github.com/snakers4/silero-models
Fairseq Transformerhttps://hf.co/facebook/ttstransformer-ru-cv7css10
SpeechT5https://hf.co/voxxer/speecht5finetunedcommonvoicerutranslit
Vosk-TTShttps://github.com/alphacep/vosk-tts
EdgeTTShttps://github.com/rany2/edge-tts
VK Cloudhttps://cloud.vk.com/
SaluteSpeechhttps://developers.sber.ru/portal/products/smartspeech
ElevenLabshttps://elevenlabs.io/