datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-bias-research-landscape
Dataset Card: AI Bias Research Landscape
Dataset Summary
This dataset contains 692 curated bibliographic records of peer-reviewed and
preprint publications on artificial intelligence (AI) and algorithmic bias,
published between 2012 and 2026. Each record includes publication metadata
(paper title, DOI, authors, author regions, affiliations, publication year,
and research domain), author ORCID identifiers, and OpenAlex-derived
metadata, including OpenAlex IDs… See the full description on the dataset page: https://huggingface.co/datasets/cair-nepal/ai-bias-research-landscape.rakshak-nepali-toxicity-augmentednepali-tts-mos-resultscomplete_nepal_share_market_datacc100-nepali-strictly-cleaned-devanagari-only
CC-100 Nepali — Cleaned(Devanagari Only)
Pipeline
Unicode normalisation (NFC + ftfy)
Rule-based filters (length, Devanagari ratio ≥ 0.5, boilerplate)
Language ID — fastText lid.176.bin, confidence ≥ 0.7
Exact deduplication (MD5)
Near-deduplication (char 13-gram bloom filter)
98/1/1 train/val/test split, seed 42
Usage
from datasets import load_dataset
ds = load_dataset("Basanta55/cc100-nepali-strictly-cleaned-devanagari-only")
rakshak-nepali-toxicity-v2nepal-cricket-ODI-matchesnepali_student_jobtype_survey
