datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nexttoken-pmkisan-domain-sft-data
NextToken pmkisan domain SFT data (v1)
Grounded multilingual QA dataset for fine-tuning
somasekhar-dev/NextToken-model-1
on the Indian government-schemes / banking-financial domain.
Generated by a pipeline (chunk source docs -> generate questions -> generate
grounded answers -> validate/assemble) using a local LLM generator, from
~57 scheme/product source documents (PM-KISAN, Ayushman Bharat, MGNREGA,
banking products, insurance, savings instruments, etc.).
Files… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-pmkisan-domain-sft-data.SomaliCrowS
SomaliCrowS: A Gender Bias Benchmark for Somali Language Models
Dataset Description
SomaliCrowS is a benchmark for measuring gender bias in Somali language models. It contains matched sentence pairs — identical except for the grammatical gender of the subject — spanning social domains where stereotyping commonly occurs, including:
Occupation
Leadership
Business
Education
STEM
Family
Politics
For each pair, a masked-language-model is queried to compute the… See the full description on the dataset page: https://huggingface.co/datasets/Abdullahicoder/SomaliCrowS.
