datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sakhi
Sakhi: A Community-Validated Multilingual Maternal-Health Benchmark
Sakhi is a benchmark for evaluating large language models on maternal and reproductive-health questions in three languages spoken in low-resource settings: English, Hindi, and Marathi. It was built around a deployed WhatsApp-based maternal-health chatbot reaching rural mothers in Hindi- and Marathi-speaking districts of India, with a three-channel review pipeline: practising Indian doctors, Accredited Social Health… See the full description on the dataset page: https://huggingface.co/datasets/SimPPL/sakhi.SA-Knowledge
SA-Knowledge
This repository collects corpora and evaluation data for four South African
languages: isiZulu, isiXhosa, Sepedi and Sesotho. The resources were developed
for the doctoral thesis Injecting Commonsense Knowledge into Pretrained
Language Models for Low Resource Languages (University of Cape Town, 2026).
Each subset corresponds to a thesis chapter and can be used independently.
Point of contact: Sello Ralethe
Supervisor: Dr. Jan Buys, Department of Computer Science… See the full description on the dataset page: https://huggingface.co/datasets/sello-ralethe/SA-Knowledge.Gong-Poem-Dataset
龚诗完整全文数据集
本仓库发布经过规范化与逐条校验的龚诗全文。当前版本为 v1.0.1-fulltext,校验日期为 2026-07-14。
数据规模
配置
文件
条目数
内容
canonical_poems
data/poems_full.parquet
17
规范诗作全文
editions
data/editions_full.parquet
24
不同来源、版本见证全文
fragments
data/fragments_full.parquet
4
可核验散句
每个配置均同时提供 Parquet、CSV 和 JSONL;Hugging Face 查看器直接读取带显式字段类型的 Parquet。所有 45 条全文记录均完成 Unicode NFC 规范化,并通过目录记录的行数、非空白字符数及 SHA-256 一致性校验;结果见 data/full_text_validation.json。
使用方式
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Sakanaction/Gong-Poem-Dataset.
