CoolFace
Datasetpublic

cidjeu/french-moore-parallel-conf-ge-0.5

French → Mooré (confidence ≥ 0.5) Subset of the full French–Mooré validated parallel corpus restricted to pairs with translation_confidence >= 0.5. Snapshot Field Value Pairs in this subset ~2.42 million Filter translation_confidence >= 0.5 Source language French Target language Mooré (Mossi) Translator Glosbe public MT Parent dataset full validated export (confidence floor ~0.35) Files fr-mos-validated-conf-ge-0.5.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/cidjeu/french-moore-parallel-conf-ge-0.5.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes10downloads
Dataset Card

French → Mooré (confidence ≥ 0.5)

Subset of the full French–Mooré validated parallel corpus restricted to pairs with `translation_confidence >= 0.5`.

Snapshot

FieldValue
Pairs in this subset~2.42 million
Filtertranslation_confidence >= 0.5
Source languageFrench
Target languageMooré (Mossi)
TranslatorGlosbe public MT
Parent datasetfull validated export (confidence floor ~0.35)

Files

  • fr-mos-validated-conf-ge-0.5.jsonl
  • fr-mos-validated-conf-ge-0.5.parquet

Schema

Same as the full corpus: id, source_domain, source_url, crawl_date, french, moore, translator, translation_confidence, complexity_score, quality_score.

Notes

  • Automatic MT; not fully human-reviewed.
  • This subset is higher-confidence than the full validated dump but still noisy.
  • Prefer this split when you want a stricter automatic quality cut at 0.5.