CoolFace
Datasetpublic

louisbertson/french-moore-parallel

French → Mooré (Mossi) Parallel Corpus Machine-translated parallel sentences from French (fr) to Mooré / Mossi (mos), produced by a public-web crawl + filtering + Glosbe translation pipeline. Snapshot Field Value Validated pairs 3,000,040 Source language French Target language Mooré (Mossi) Translator Glosbe public MT Export date 2026-08-14 Schema Column Type Description id string (UUID) Pair identifier… See the full description on the dataset page: https://huggingface.co/datasets/louisbertson/french-moore-parallel.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
1likes67downloads
Dataset Card

French → Mooré (Mossi) Parallel Corpus

Machine-translated parallel sentences from French (fr) to Mooré / Mossi (mos), produced by a public-web crawl + filtering + Glosbe translation pipeline.

Snapshot

FieldValue
Validated pairs3,000,040
Source languageFrench
Target languageMooré (Mossi)
TranslatorGlosbe public MT
Export date2026-08-14

Schema

ColumnTypeDescription
idstring (UUID)Pair identifier
source_domainstringCrawl source domain
source_urlstringOriginal page URL
crawl_datestring (ISO 8601)When the page was crawled
frenchstringSource sentence
moorestringMooré translation
translatorstringProvider id (e.g. glosbe)
translation_confidencefloatProvider confidence
complexity_scorefloatSource complexity score
quality_scorefloatPipeline quality score

Files

  • fr-mos-validated.jsonl — one JSON object per line
  • fr-mos-validated.parquet — same rows, ZSTD-compressed Parquet

Quality notes

  • Only accepted / validated pairs are included (failed validation rows are excluded).
  • Translations are automatic (Glosbe); this is not a fully human-reviewed gold set.
  • Source text was filtered for short/medium simple French sentences; exact duplicates were deduplicated.
  • Soft confidence floor was applied during generation; average confidence on accepted pairs is roughly ~0.63.
  • Dominant crawl sources include French Wikipedia and other public French sites; review licenses/terms for your use case.

Intended use

  • Pretraining / fine-tuning low-resource MT models for French ↔ Mooré
  • Research on West African language technology
  • Bootstrapping human post-editing workflows

Limitations

  • Automatic MT noise and occasional domain skew (encyclopedia, news, institutional)
  • Not a substitute for native-speaker verified data for high-stakes applications
  • Source provenance is retained in source_url / source_domain for auditability

Citation

If you use this dataset, please cite the dataset page and note that translations were generated with Glosbe MT via an automated corpus pipeline.