datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
talent_plus_rl_groups_of_50_with_audiobox_scoresstrudel-rl-data
strudel-rl-data
Datasets from the Strudel-RL project (training a Qwen3.6-35B-A3B to write house / techno / hypnotic techno as Strudel programs).
sft/: SFT datasets v1..v5 (messages format, system+user+assistant) with stats.
synth/: GLM-5.3-Flash synthetic rounds 1..5 (prompt, code, reasoning length).
captions/: GLM captions of programs. prompts/: brief generators and held-out eval briefs (v0..v3; v3 = reference-driven).
gold/: 26 hand-written gold programs + manifest.… See the full description on the dataset page: https://huggingface.co/datasets/amol-derick/strudel-rl-data.SoundMindDataset
SoundMind Dataset
SoundMind: RL-Incentivized Logic Reasoning for Audio-Language Models (EMNLP 2025)
SoundMind Dataset is an Audio Logical Reasoning (ALR) dataset consisting of 6,446 text-audio annotated samples specifically designed for complex reasoning tasks, where models determine whether conclusions are "entailed" or "not-entailed" based on logical premises, featuring comprehensive chain-of-thought reasoning annotations in both audio and text formats as the first audio-level… See the full description on the dataset page: https://huggingface.co/datasets/SoundMind-RL/SoundMindDataset.new-rl-vitaReagent-RL-709K
Official Repo of Reagent Agent RL training dataset (Reagent-RL-709K).
Paper: https://arxiv.org/abs/2601.22154
Abstract:
Agentic Reinforcement Learning (Agentic RL) has achieved notable success in enabling agents to perform complex reasoning and tool use.
However, most methods still relies on sparse outcome-based reward for training.
Such feedback fails to differentiate intermediate reasoning quality, leading to suboptimal training results.
In this paper, we introduce… See the full description on the dataset page: https://huggingface.co/datasets/bunny127/Reagent-RL-709K.kencorpus_sw_culture
KenCorpus Swahili Culture Subset
A filtered subset of Kencorpus/KenCorpus_audio,
containing only rows where language=Swahili and genre=Culture (37 clips).
Audio files are in audio/, indexed by kencorpus_sw_culture.jsonl with path and duration fields,
following the layout of kyutai/DailyTalkContiguous.
swa_lug_tts
Luganda-Swahili Speaker-Clustered TTS Dataset
Dataset Summary
A cleaned, speaker-labeled text-to-speech (TTS) dataset covering Luganda and Kiswahili, derived from the Luganda-Swahili Speech for Text-to-Speech Synthesis Kaggle dataset. The source data contains recordings from 6 speakers per language, but speaker identity was not labeled in the original release. Speaker labels in this version were recovered via unsupervised audio clustering (6 clusters per language)… See the full description on the dataset page: https://huggingface.co/datasets/rlabz/swa_lug_tts.mwanamke_moshi
Swahili Moshi Fine-tuning Dataset (mwanamke)
Overview
This dataset prepares Swahili conversational audio for fine-tuning
Moshika (the female-voice
Moshi variant) using
kyutai-labs/moshi-finetune.
It builds on the stereo, speaker-separated audio chunks from
rlabz/qsuperposition_mwanamke
and adds the .jsonl index and per-file .json transcripts that
moshi-finetune requires for training.
Source Data
Origin: rlabz/qsuperposition_mwanamke — stereo… See the full description on the dataset page: https://huggingface.co/datasets/rlabz/mwanamke_moshi.qsuperposition_mwanamke
Swahili Culture Conversational Dataset (Stereo)
Overview
This dataset contains conversational Swahili audio, restructured into
diarized, speaker-separated stereo chunks suitable for fine-tuning
speech-to-speech dialogue models — specifically prepared for fine-tuning
Moshika (the female-voice Moshi variant) via
kyutai-labs/moshi-finetune.
The processing pipeline was adapted from the approach used in
rlabz/sample_sw_culture.
Source Data
Origin:… See the full description on the dataset page: https://huggingface.co/datasets/rlabz/qsuperposition_mwanamke.KenSpeech
KenSpeech: A Swahili Speech Dataset for ASR
Dataset Description
KenSpeech is a comprehensive Swahili speech dataset containing both read and spontaneous speech recordings from native Swahili speakers in Kenya. This dataset is designed for training and evaluating automatic speech recognition (ASR) and speech-to-text (STT) systems for Swahili.
Dataset Statistics
Metric
Value
Total Duration
27 hours 31 minutes 50 seconds
Read Speech… See the full description on the dataset page: https://huggingface.co/datasets/rlox/KenSpeech.sample_sw_culture
Swahili Culture Conversational Dataset (Stereo)
Overview
This dataset contains conversational Swahili audio samples derived from the
Culture genre subset of KenCorpus_audio
(CC-BY-4.0), restructured into diarized, speaker-separated stereo audio
chunks suitable for fine-tuning speech-to-speech dialogue models such as
Moshi / Hibiki.
Processing notebook: the full pipeline (diarization, conversational
chunking, stereo construction, and upload) is available here:… See the full description on the dataset page: https://huggingface.co/datasets/rlabz/sample_sw_culture.original_1900_1999original_1500_1599vmar-rl
VMAR — RL Prompt Set (GRPO / RLVR)
Audio-reasoning prompt set for verifiable-reward RL (verl GRPO). No assistant target —
each row is a prompt + locked gold (answer, per-hop gold_spans) for a programmatic
reward (exact-match outcome + per-hop citation tIoU).
Stats
train 69,534 · val 8,828 · test 8,925 · ~3,061 audio-hours
modality sound/music/speech · n_hops 2–30 (every clip >30s → multi-chunk)
Configs
default (train/validation/test): id… See the full description on the dataset page: https://huggingface.co/datasets/ngqtrung/vmar-rl.working1new-rlMeeami_RLHFexpressive-eng-tts
r-labs/expressive-eng-tts
Expressive synthetic Ugandan English speech dataset for conversational Text-to-Speech (TTS) fine-tuning.
Dataset Summary
r-labs/expressive-eng-tts is a fully synthetic expressive Ugandan English TTS dataset designed for fine-tuning conversational speech models with authentic Ugandan English accent, prosody, and expressive speaking behaviors.
The dataset contains speech generated from 3 synthetic speakers:
2 Female speakers
1 Male… See the full description on the dataset page: https://huggingface.co/datasets/r-labs/expressive-eng-tts.joy_sw-v0original_1600_1699original_1700_1799original_1800_1899MEEAMI_RLHF_PILOTjoy_sw-v0-bilingualsalt-eng-tts-inference-v0telugu_rl
