datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
french-moore-parallel
French → Mooré (Mossi) Parallel Corpus
Machine-translated parallel sentences from French (fr) to Mooré / Mossi (mos), produced by a public-web crawl + filtering + Glosbe translation pipeline.
Snapshot
Field
Value
Validated pairs
3,000,040
Source language
French
Target language
Mooré (Mossi)
Translator
Glosbe public MT
Export date
2026-08-14
Schema
Column
Type
Description
id
string (UUID)
Pair identifier… See the full description on the dataset page: https://huggingface.co/datasets/louisbertson/french-moore-parallel.french-moore-parallel-conf-ge-0.5
French → Mooré (confidence ≥ 0.5)
Subset of the full French–Mooré validated parallel corpus restricted to pairs with
translation_confidence >= 0.5.
Snapshot
Field
Value
Pairs in this subset
~2.42 million
Filter
translation_confidence >= 0.5
Source language
French
Target language
Mooré (Mossi)
Translator
Glosbe public MT
Parent dataset
full validated export (confidence floor ~0.35)
Files
fr-mos-validated-conf-ge-0.5.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/louisbertson/french-moore-parallel-conf-ge-0.5.french-moore-parallel
French → Mooré (Mossi) Parallel Corpus
Machine-translated parallel sentences from French (fr) to Mooré / Mossi (mos), produced by a public-web crawl + filtering + Glosbe translation pipeline.
Snapshot
Field
Value
Validated pairs
3,000,040
Source language
French
Target language
Mooré (Mossi)
Translator
Glosbe public MT
Export date
2026-08-14
Schema
Column
Type
Description
id
string (UUID)
Pair identifier… See the full description on the dataset page: https://huggingface.co/datasets/cidjeu/french-moore-parallel.long-nature-21beb2
long-nature-21beb2
Synthetic weather test data: 42 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/moorelaura49/long-nature-21beb2.moore-audio-standardized ---
pretty_name: louisbertson/moore-audio-standardized
language:
- mos
tags:
- audio
- moore
- self-supervised-learning
- speech
size_categories:
- n<1K
---
# Mooré Standardized Audio Dataset
This dataset was exported from the preprocessing pipeline in this repository. It keeps the repository's canonical split manifests and uses standardized WAV audio so the same files work in local training, Google Colab, and Hugging Face Hub uploads.… See the full description on the dataset page: https://huggingface.co/datasets/louisbertson/moore-audio-standardized.small-week-71a026
small-week-71a026
Synthetic products test data: 50 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/James-Moore/small-week-71a026.so101_task_putinto_20260529_150522This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/MooreFoss/so101_task_putinto_20260529_150522.so101_task_putinto_20260529_151343This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/MooreFoss/so101_task_putinto_20260529_151343.french-moore-parallel-conf-ge-0.5
French → Mooré (confidence ≥ 0.5)
Subset of the full French–Mooré validated parallel corpus restricted to pairs with
translation_confidence >= 0.5.
Snapshot
Field
Value
Pairs in this subset
~2.42 million
Filter
translation_confidence >= 0.5
Source language
French
Target language
Mooré (Mossi)
Translator
Glosbe public MT
Parent dataset
full validated export (confidence floor ~0.35)
Files
fr-mos-validated-conf-ge-0.5.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/cidjeu/french-moore-parallel-conf-ge-0.5.moore-fr-momoore-flores
