datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VCDB-Core-AudioVideo
VCDB Core Audio-Video Retrieval
This repository packages synchronized video and extracted audio from the
528-video core set of VCDB as a symmetric video+audio-to-video+audio
retrieval task for MTEB/MOEB. The separate 100,000-video background collection
is not included.
Terms and provenance
The source dataset is provided by Fudan University for research purposes
only. The source authors and Fudan University make no warranties about the
dataset, including… See the full description on the dataset page: https://huggingface.co/datasets/pranitchawla/VCDB-Core-AudioVideo.VCDB-Core-Audio
VCDB Core Audio Retrieval
This repository packages audio extracted from the 528-video core set of
VCDB as a symmetric audio-to-audio retrieval task for MTEB/MOEB. The separate
100,000-video background collection is not included.
Terms and provenance
The source dataset is provided by Fudan University for research purposes
only. The source authors and Fudan University make no warranties about the
dataset, including non-infringement. Users must review and follow the… See the full description on the dataset page: https://huggingface.co/datasets/pranitchawla/VCDB-Core-Audio.cores_es-ljspeech
Cores — Español (es)
LJSpeech dataset of Cores (es).
327 pairs
44.1kHz mono 16-bit PCM WAV (original wiki quality)
Piper TTS Training (High Quality on T4 GPU)
Preprocessing (downsample to 22.05kHz)
python3 -m piper_train.preprocess \
--language es \
--input-dir ./cores_es \
--output-dir ./train_cores_es \
--dataset-format ljspeech \
--single-speaker \
--sample-rate 22050
Training (Kaggle T4 16GB)
python3 -m piper_train… See the full description on the dataset page: https://huggingface.co/datasets/RoxasYTB/cores_es-ljspeech.coretocores_de-ljspeech
Cores — Deutsch (de)
LJSpeech dataset of Cores (de).
269 pairs
44.1kHz mono 16-bit PCM WAV (original wiki quality)
Piper TTS Training (High Quality on T4 GPU)
Preprocessing (downsample to 22.05kHz)
python3 -m piper_train.preprocess \
--language de \
--input-dir ./cores_de \
--output-dir ./train_cores_de \
--dataset-format ljspeech \
--single-speaker \
--sample-rate 22050
Training (Kaggle T4 16GB)
python3 -m piper_train… See the full description on the dataset page: https://huggingface.co/datasets/RoxasYTB/cores_de-ljspeech.cores_ru-ljspeech
Cores — Русский (ru)
LJSpeech dataset of Cores (ru).
327 pairs
44.1kHz mono 16-bit PCM WAV (original wiki quality)
Piper TTS Training (High Quality on T4 GPU)
Preprocessing (downsample to 22.05kHz)
python3 -m piper_train.preprocess \
--language ru \
--input-dir ./cores_ru \
--output-dir ./train_cores_ru \
--dataset-format ljspeech \
--single-speaker \
--sample-rate 22050
Training (Kaggle T4 16GB)
python3 -m piper_train… See the full description on the dataset page: https://huggingface.co/datasets/RoxasYTB/cores_ru-ljspeech.CoReBench_v1
Dataset Card for CoReBench_v1
Dataset Summary
COREBench is a comprehensive conversational reasoning benchmark designed to evaluate audio language models on reasoning capabilities in multi-turn conversations.
Example Instance
Question: What is the fruit the first speaker likes most?
Audio Sample:
Download audio: https://huggingface.co/datasets/chiheemwong/CoReBench_v1/audio/ebd9de53fbca567cf675.mp3
[Transcript]
Zinaida: Alright team, let's nail this chorus. We… See the full description on the dataset page: https://huggingface.co/datasets/chiheemwong/CoReBench_v1.cores_en-ljspeech
Cores — English (en)
LJSpeech dataset of Cores (en).
381 pairs
44.1kHz mono 16-bit PCM WAV (original wiki quality)
Piper TTS Training (High Quality on T4 GPU)
Preprocessing (downsample to 22.05kHz)
python3 -m piper_train.preprocess \
--language en-us \
--input-dir ./cores_en \
--output-dir ./train_cores_en \
--dataset-format ljspeech \
--single-speaker \
--sample-rate 22050
Training (Kaggle T4 16GB)
python3 -m… See the full description on the dataset page: https://huggingface.co/datasets/RoxasYTB/cores_en-ljspeech.cores_fr-ljspeech
Cores — Français (fr)
LJSpeech dataset of Cores (fr).
327 pairs
44.1kHz mono 16-bit PCM WAV (original wiki quality)
Piper TTS Training (High Quality on T4 GPU)
Preprocessing (downsample to 22.05kHz)
python3 -m piper_train.preprocess \
--language fr \
--input-dir ./cores_fr \
--output-dir ./train_cores_fr \
--dataset-format ljspeech \
--single-speaker \
--sample-rate 22050
Training (Kaggle T4 16GB)
python3 -m piper_train… See the full description on the dataset page: https://huggingface.co/datasets/RoxasYTB/cores_fr-ljspeech.
