lblommesteyn/papuan-climate-science-corpus
Papuan Climate Science Corpus Dataset Summary This dataset contains 500 records documenting climate science and environmental knowledge in two severely under-documented Papuan languages from Papua New Guinea, the most linguistically diverse country in the world. Covered Languages Language ISO Code Speakers Status Region Usan wnu ~1,400 Vulnerable Madang Province, PNG Yélî Dnye yle ~3,000 Vulnerable Rossel Island, PNG… See the full description on the dataset page: https://huggingface.co/datasets/lblommesteyn/papuan-climate-science-corpus.
Papuan Climate Science Corpus
Dataset Summary
This dataset contains 500 records documenting climate science and environmental knowledge in two severely under-documented Papuan languages from Papua New Guinea, the most linguistically diverse country in the world.
Covered Languages
Why This Dataset Matters
- Papua New Guinea has 800+ languages (12% of world's total)
- Most Papuan languages have zero digital documentation
- Climate knowledge in indigenous languages is disappearing as languages become endangered
- Usan and Yélî Dnye are among the least documented languages globally
Dataset Structure
Use Cases
- Indigenous climate knowledge preservation
- Climate adaptation research in Papua New Guinea
- Computational linguistics for severely under-resourced languages
- Environmental policy development with local communities
License
CC0 1.0 Universal (Public Domain Dedication)
