flags
Datasets
All datasets matching “flags”savtadepth-flags-V2cameroon_bibles
Cameroon Bibles — verse-aligned scripture corpus
The text corpus behind Lingo / NativeAI: verse-aligned scripture
across 60 Cameroonian languages (64 translation versions). Scripture is one of
the few sources of sentence-aligned parallel text for these low-resource languages — the
aligned backbone of our corpus (see the research log).
Layout
<Language>/<BOOK>.<chapter>.txt e.g. Ngi/MAT.2.txt
Each file is one chapter; lines are verse-numbered, alignable across… See the full description on the dataset page: https://huggingface.co/datasets/flagship-ai/cameroon_bibles.Rick-bot-flags
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/kingabzpro/Rick-bot-flags.Urdu-ASR-flags
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/kingabzpro/Urdu-ASR-flags.savtadepth-flagsghomala-spoken-bible
Ghomálá' Spoken New Testament — aligned audio + trilingual text
Part of the Lingo / NativeAI language-preservation project. This is
~20 hours of spoken Ghomálá' (Ghomala, ISO bbj; a Grassfields Bantu language of
West Cameroon) — recorded readings of the New Testament — aligned chapter-by-chapter
with parallel text in Ghomálá', French, and English.
Spoken-language data is exactly what oral-first Cameroonian languages lack, which makes
this a rare resource for building ASR, TTS… See the full description on the dataset page: https://huggingface.co/datasets/flagship-ai/ghomala-spoken-bible.
