datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
distilled-web
Dataset Card for lenamerkli/distilled-web
This dataset consists of web-scraped data using a custom crawler purpose-built for each website.
Dataset Details
Dataset Sources
Repository: https://github.com/lenamerkli/distilled-web
Uses
This dataset is useful for training large language models.
The train split provides instruction-following and chat data for supervised fine-tuning (SFT) and instruction tuning.
The pretrain split… See the full description on the dataset page: https://huggingface.co/datasets/lenamerkli/distilled-web.HeySQuAD_distilldistill-whisper-fin-jargonpersonaplex-distill-conversations
PersonaPlex Distillation Conversation Dataset
Teacher-generated multi-turn conversation data for distilling/pruning NVIDIA PersonaPlex 7B
(a Moshi-architecture full-duplex speech-to-speech model).
What this is
Each sample is a real conversation rendered by the PersonaPlex teacher itself:
The student's turns are scripted (generated by Qwen3-8B) and voiced with Piper TTS
The teacher's (PersonaPlex's) responses are improvised live by the model — its real… See the full description on the dataset page: https://huggingface.co/datasets/niloy629/personaplex-distill-conversations.modded-distill-wavlm-base
Dataset Summary
Lance tables of LibriSpeech utterances at 16 kHz with cached last-layer representations from frozen microsoft/wavlm-base.
This release is a precomputed teacher cache plus raw audio bytes, not a general-purpose speech benchmark split.
Structure
Local / mirrored Hub layout:
Directory
Content
train/
LibriSpeech 960 h train corpora: train-clean-100, train-clean-360, train-other-500.
eval/
LibriSpeech dev-clean.
Each of train/ and eval/ is a… See the full description on the dataset page: https://huggingface.co/datasets/Alright7398/modded-distill-wavlm-base.vox2_3D_distill_shard22vox2_3D_distill_shard24vox2_3D_distill_shard21vox2_3D_distill_shard23vox2_3D_distill_shard13vox2_3D_distill_shard17vox2_3D_distill_shard16vox2_3D_distill_shard40vox2_3D_distill_shard42vox2_3D_distill_shard14vox2_3D_distill_shard32distill-4vox2_3D_distill_shard20maya-distill-datamultilingual-test-distill-strong-tts-20260520
Multilingual Test Distill Strong TTS 20260520
This repository contains a distributable tar-sharded version of multilingual_test_distill_strong_tts_20260520.
The dataset follows the local voice_dataset/data layout after extraction:
data/csvs/metadata_zh.csv
data/csvs/metadata_en.csv
data/csvs/metadata_ja.csv
data/csvs/metadata_ko.csv
data/zh/**/*.wav
data/en/**/*.wav
data/ja/**/*.wav
data/ko/**/*.wav
Metadata format:
file_path|duration|dnsmos|text
dnsmos is intentionally blank… See the full description on the dataset page: https://huggingface.co/datasets/guangzhaoli/multilingual-test-distill-strong-tts-20260520.chunk_20chunk_204chunk_237chunk_208nepali-oov-distilled
Nepali OOV-distilled subset (854 h)
An OOV-dense distillation of Premal-12/c9nepali-audio-dataset2 (used with the
author's permission), shipped in four variants: the original single-voice audio,
a CPU-augmented copy, and 244 h re-rendered onto 1,842 real human speakers with
Seed-VC. For Nepali ASR and TTS work.
Filter with the variant field -- see Composition below. If you came here
for speaker diversity, you want variant == "vc".
What this is
The source corpus is… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-oov-distilled.chunk_19chunk_38chunk_142chunk_209vox2_3D_distill_shard_06
