datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cv22_azeros
FLEURS (Lhotse cuts)
Each language is a separate config. Load a single language's cuts as a HF Dataset of raw manifest records with, e.g.:
from datasets import load_dataset
ds = load_dataset("your-org/REPO_NAME", "bg_bg", split="train")
If audio shards (recording.NNNNN.tar) are present alongside the cuts, the LANG/SPLIT/ folder is a valid Lhotse Shar directory. Download it (e.g. via snapshot_download) and load with Lhotse directly:
from huggingface_hub import snapshot_download… See the full description on the dataset page: https://huggingface.co/datasets/sonalsannigrahi/cv22_azeros.AudioSkills-Llama3
AudioSkills-XL Dataset
To promote the development of open source models, we have released AudioSkills using the exact same method generated with Llama 3.1-8B Instruct instead of GPT4o in the original.
Project page | Paper | Code
Dataset Description
AudioSkills-XL is a large-scale audio question-answering (AQA) dataset designed to develop (large) audio-language models on expert-level reasoning and problem-solving tasks over short audio clips (≤30 seconds). It… See the full description on the dataset page: https://huggingface.co/datasets/sonalkum/AudioSkills-Llama3.RealWorldQuestioning
RealWorldQuestioning Benchmark
RealWorldQuestioning is a benchmark dataset of 400+ real-world user questions collected from public discussion forums (e.g., Reddit, Quora), designed to support evaluation of gender bias and information disparity in Large Language Models (LLMs). The dataset spans four business-relevant domains: Education, Jobs, Investment, and Health.
Each question is annotated with:
User persona (Male or Female framing)
Source forum
Domain category
Four anonymized… See the full description on the dataset page: https://huggingface.co/datasets/SonalPrabhune/RealWorldQuestioning.MedExpert
MedExpert
MedExpert-Benchmark features clinician-created questions and detailed annotations designed to assess the accuracy, completeness, and reliability of LLM-generated medical responses.
It comprises 540 question–response pairs across two distinct specialties:
Young Adult Mental Health (MH)
Prenatal Care (PC)
Each sample is annotated by clinical subject-matter experts for factual accuracy, completeness (omissions), and model certainty. This dataset is designed to support… See the full description on the dataset page: https://huggingface.co/datasets/sonal-ssj/MedExpert.sona-corpus
THE SONA CORPUS — Noisy-to-Clean Hindi–English Parallel Dataset
A clean, bilingual dataset card you can read at a glance and use immediately.
Curated by: Aditya (AADIMIND)
Languages: Hindi, English
Total examples: 581312 (INPUT: 256 TOKEN• TARGET: 256 TOKEN)
Tasks: Text cleaning, GEC, OCR post-processing, Seq2Seq fine-tuning
License: MIT
Source: Hindi Wikipedia (HiWiki) processed into noisy–clean pairs
Repo: https://huggingface.co/datasets/AADIMIND/sona-corpus… See the full description on the dataset page: https://huggingface.co/datasets/AADIMIND/sona-corpus.LoFTI
LoFTI: Localization and Factuality Transfer to Indian Locales
LoFTI is a benchmark dataset that can be used to evaluate an LLM’s localization and factual text transfer capabilities. LoFTI consists of factual statements about entities in source and target locations; the source locations are spread across the globe and the target locations are all within India with varying degrees of hyperlocality (country, states, cities). The entities span a wide variety of categories.… See the full description on the dataset page: https://huggingface.co/datasets/sonasimon/LoFTI.cakewalk_sonar_x2_god_producer
Cakewalk Sonar X2 — God-Level Producer Dataset (8k)
Train an LLM to become a god-level music producer in Cakewalk Sonar X2 Producer Edition.
This dataset contains 4,060 high-quality examples covering every aspect of professional music production in Sonar X2:
Virtual instrument programming (Dimension Pro, Rapture, Z3TA+, etc.)
Plugin mastery (ProChannel, Sonitus suite, Console Emulator, etc.)
Sound design & synthesis
Advanced editing (AudioSnap, VocalSync, Step Sequencer)
Mixing… See the full description on the dataset page: https://huggingface.co/datasets/11-47/cakewalk_sonar_x2_god_producer.sonar-municipal-pl-actions
Sonar Municipal — PL Actions Corpus
The largest publicly released Portuguese legal text rewriting dataset: 241,111 pairs mapping the original ementa (summary) of a Brazilian municipal Projeto de Lei (PL) to its action-form textualization (ação). Action form is a direct, imperative rewrite that strips juridical boilerplate, normalizes typography, and surfaces the underlying intervention rather than the legal instrument.
Companion to: ICMC-USP undergraduate thesis "Mineração… See the full description on the dataset page: https://huggingface.co/datasets/thiagoambiel/sonar-municipal-pl-actions.cakewalk_sonar_8_god_producer_dataset
Cakewalk Sonar 8 Producer Edition - God Level Producer Dataset
The ultimate training dataset for becoming a god-level producer in Cakewalk Sonar 8 Producer Edition.
This dataset trains LLMs to master every aspect of Sonar 8 Producer Edition at the highest professional level:
All virtual instruments (Dimension Pro, Rapture, Pentagon I, etc.)
ProChannel (Console Emulator, EQ, Compressor, Tube Saturation)
Complete mixing workflows
Commercial mastering chains (LP-64, Sonitus:fx)… See the full description on the dataset page: https://huggingface.co/datasets/11-47/cakewalk_sonar_8_god_producer_dataset.mls_enml-intern-llama-sft-data
ML-Intern Llama SFT data
Prepared local session logs for TRL SFTTrainer. Rows use conversational prompt/completion format.
sonamdatamls_ptyodas_sampleVIDHI-Indian_Law_Consumer_Enrichedjtdatasetinstruction_dataset-edu-aimarvel-characters-dataset
