ghananlpcommunity/ghana-corpus
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Ghana Corpus Verse-aligned text for Ghanaian languages, plus several world languages, for building monolingual and parallel corpora. Every language is aligned on a shared verse key, so any single language can be pulled on its own or… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-corpus.
Standardize license to CC BY-NC 4.0 and add research-use notice
Add 8 new corpus file(s)
Add 2 new corpus file(s)
Add 2 new corpus file(s)
Add 3 new corpus file(s)
Add 1 new corpus file(s)
Reframe as monolingual-and-parallel
Document multiple reference versions + @version selection
Add contemporary reference Bible versions (9 languages)
Sync Ghana corpus data
Sync Ghana corpus data
Add 4 new corpus file(s)
Add 1 new corpus file(s)
Add 2 new corpus file(s)
Add 14 new corpus file(s)
Add 5 new corpus file(s)
Reference files now self-describing via filename
Add 8 new corpus file(s)
Remove old code-only reference filenames
Document version_id/lang_code mapping for reference config
Sync Ghana corpus data
Define ghanaian/english/reference configs to fix dataset viewer
Remove end-to-end test file
Add 1 new corpus file(s)
Point dataset card to GhanaNLP repo
Add dataset card pointing to the GitHub repo
Update Ghana corpus data
initial commit
