datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CrawlSinger-OS
CrawlSinger-OS
CrawlSinger-OS is a large-scale, open-source singing corpus constructed for
score-native singing voice synthesis. It contains more than 2,300 hours of
processed singing data from multiple public song and singing collections, with
a unified annotation scheme for lyrics, MIDI pitches, symbolic note values,
lyric-to-note alignment, and global tempo.
VocalRender paper
VocalRender code
VocalRender checkpoints
Why CrawlSinger-OS
Modern singing… See the full description on the dataset page: https://huggingface.co/datasets/pymaster/CrawlSinger-OS.web-crawl-v1
OpenTransformers Web Crawl v1
Your data. Your company. No apologies.
Stats
Total pages: 45,026
Total text: 651.3 MB
Crawled: 2026-01-13
Format
JSONL (gzipped), one document per line:
{
"url": "https://example.com/page",
"domain": "example.com",
"timestamp": "2026-01-13T02:43:19.685727",
"status": 200,
"text": "Clean extracted text content...",
"text_len": 1234,
"html_len": 5678,
"links": 42,
"fetch_ms": 150,
"hash": "abc123..."
}… See the full description on the dataset page: https://huggingface.co/datasets/OpenTransformer/web-crawl-v1.crawl_reaction_videomid-accent-crawl-youtubeyoutube_crawl_audio_jaASR_Data_Crawl_Youtubeyoutube_crawl_audio_tscribed_audioyoutube_crawl_audio_comp_audiotts-crawl
