datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
moore_audio_data~70h of Mooré paired audio+text data from jw.org using jwsoup
sampling rate = 24 KHz
Headers: audio, text
moore_audio_data~70h of Mooré paired audio+text data from jw.org and jwsoup
Headers: audio, text
tts-moore-hommeStateBench
StateBench
StateBench is a benchmark for world-state reasoning in video continuation, introduced in the paper Do Video Generators Track the World Across Segments? A Benchmark and Method for World-State Reasoning in Video Continuation.
Code: https://github.com/AMAP-ML/StateAgent
The benchmark contains 200 cross-segment continuation tasks across three difficulty levels:
Level
Count
Description
past_visible
85
The target object was visible in a prior frame — tests… See the full description on the dataset page: https://huggingface.co/datasets/moore12138/StateBench.moore-speech-bibleEVR_datasetfrench-moore-parallel
French → Mooré (Mossi) Parallel Corpus
Machine-translated parallel sentences from French (fr) to Mooré / Mossi (mos), produced by a public-web crawl + filtering + Glosbe translation pipeline.
Snapshot
Field
Value
Validated pairs
3,000,040
Source language
French
Target language
Mooré (Mossi)
Translator
Glosbe public MT
Export date
2026-08-14
Schema
Column
Type
Description
id
string (UUID)
Pair identifier… See the full description on the dataset page: https://huggingface.co/datasets/louisbertson/french-moore-parallel.french-moore-parallel-conf-ge-0.5
French → Mooré (confidence ≥ 0.5)
Subset of the full French–Mooré validated parallel corpus restricted to pairs with
translation_confidence >= 0.5.
Snapshot
Field
Value
Pairs in this subset
~2.42 million
Filter
translation_confidence >= 0.5
Source language
French
Target language
Mooré (Mossi)
Translator
Glosbe public MT
Parent dataset
full validated export (confidence floor ~0.35)
Files
fr-mos-validated-conf-ge-0.5.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/louisbertson/french-moore-parallel-conf-ge-0.5.french-moore-parallel
French → Mooré (Mossi) Parallel Corpus
Machine-translated parallel sentences from French (fr) to Mooré / Mossi (mos), produced by a public-web crawl + filtering + Glosbe translation pipeline.
Snapshot
Field
Value
Validated pairs
3,000,040
Source language
French
Target language
Mooré (Mossi)
Translator
Glosbe public MT
Export date
2026-08-14
Schema
Column
Type
Description
id
string (UUID)
Pair identifier… See the full description on the dataset page: https://huggingface.co/datasets/cidjeu/french-moore-parallel.chatbot
clean.py
Dataset Summary
A security dataset with audio text modality, stored in tfrecord format.
Preprocessing & Augmentation
Preprocessing: standard
Augmentation: autoaugment
Splits & Sampling
Split strategy: leave one out
Sampling: random
Quality & Labeling
Quality filtering: strict
Labeling: pseudo label
Files
clean.py — main artifact of this repository
License
See the license… See the full description on the dataset page: https://huggingface.co/datasets/Moore-samuel/chatbot.long-nature-21beb2
long-nature-21beb2
Synthetic weather test data: 42 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/moorelaura49/long-nature-21beb2.moore-audio-standardized ---
pretty_name: louisbertson/moore-audio-standardized
language:
- mos
tags:
- audio
- moore
- self-supervised-learning
- speech
size_categories:
- n<1K
---
# Mooré Standardized Audio Dataset
This dataset was exported from the preprocessing pipeline in this repository. It keeps the repository's canonical split manifests and uses standardized WAV audio so the same files work in local training, Google Colab, and Hugging Face Hub uploads.… See the full description on the dataset page: https://huggingface.co/datasets/louisbertson/moore-audio-standardized.small-week-71a026
small-week-71a026
Synthetic products test data: 50 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/James-Moore/small-week-71a026.moore-aikinggroup-proverbesmoore_audio_bible
Total duration = ~24 hours scraped from bible.com
sampling rate = 24 KHz
custom_llm_001so101_task_putinto_20260529_150522This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/MooreFoss/so101_task_putinto_20260529_150522.yaabavoice-moore-tts-v0moore-instruct-stage2moore-fr-translationnew-moore-speech-clean
Moore Speech Proverbs: A Parallel Audio-Text Corpus for Mooré and French
The Moore Speech Proverbs dataset is a bilingual audio-text corpus of traditional proverbs in Mooré and French, designed for research and academic purposes in low-resource speech and language processing.
It is intended primarily for academic or research purposes in text-to-speech (TTS) and automatic speech recognition (ASR) for Mooré language.
[!NOTE]
⚠️ Access is gated. To request access, please read the… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/new-moore-speech-clean.english-moore_sentence-pairs_mt560
English-Moore Parallel Dataset
This dataset contains parallel sentences in English and Moore (Burkina Faso).
Dataset Information
Language Pair: English ↔ Moore
Language Code: mos
Country: Burkina Faso
Original Source: OPUS MT560 Dataset
Dataset Structure
The dataset contains parallel sentences that can be used for:
Machine translation training
Cross-lingual NLP tasks
Language model fine-tuning
Citation
If you use this dataset, please cite the… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-moore_sentence-pairs_mt560.ctfgujhcyso101_task_putinto_20260529_151343This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/MooreFoss/so101_task_putinto_20260529_151343.fr-moore-reasoning-v1
FR-Mooré Reasoning v1
Bilingual French ↔ Mooré (mos) dataset pairing each translation with a
reasoning trace in French that explains, structurally, how one side maps to
the other. Built by BurkimbIA to train
reasoning-capable translation models for Mooré, a low-resource language of
Burkina Faso.
Dataset at a glance
Records
248 766 (236 327 train / 12 439 test)
Source pairs
124 383, each expanded into both directions (fr→moore, moore→fr)
Source… See the full description on the dataset page: https://huggingface.co/datasets/burkimbia/fr-moore-reasoning-v1.huberman_lab_LIVE_EVENT_QA_Dr__Andrew_Huberman_at_the_Moore_Theatre_in_Seattlegoai-moore-speech-contes
Moore Speech Contes: A Spoken Corpus of Traditional Mooré Stories
The Moore Speech Contes dataset is a collection of spoken folk stories (contes) in Mooré, designed for research and academic purposes in low-resource speech and language processing.
It is intended primarily for academic or research purposes in text-to-speech (TTS) and automatic speech recognition (ASR) for Mooré language (ISO 639-3: mos).
[!NOTE]
⚠️ Access is gated. To request access, please read the policy below.🚩… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/goai-moore-speech-contes.quran-audio-moore
Dataset Coran en Mooré
Dataset d'enregistrements audio du Coran traduit en langue Mooré (Burkina Faso).
Structure
data/ : Fichiers audio WAV 16kHz
metadata.csv : Métadonnées des enregistrements
Format des données
Chaque enregistrement contient :
L'audio du verset en mooré
L'identifiant du verset
Le numéro de sourate et de verset
Le texte traduit en mooré
Les informations sur le contributeur
La date d'enregistrement
Utilisation
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/sheickydollar/quran-audio-moore.moore-speech-devinettes
Moore Speech Devinettes: A Spoken Riddle Dataset in Mooré
The Moore Speech Devinettes dataset is a spoken collection of traditional Mooré riddles, created for academic and research use in low-resource speech and language technologies.
It is designed to support work in text-to-speech (TTS), automatic speech recognition (ASR), and oral tradition modeling for the Mooré language (ISO 639-3: mos).
[!NOTE]
⚠️ Access is gated. To request access, please read the policy below.
🚩 TLDR: For… See the full description on the dataset page: https://huggingface.co/datasets/anyantudre/moore-speech-devinettes.so101_task_putinto
