datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Ganit
Ganit: A Difficulty-Aware Bengali Mathematical Reasoning Dataset
Dataset Description
Ganit (গণিত, Bengali for "mathematics") is a rigorously-processed, difficulty-aware Bengali mathematical reasoning dataset designed for training and evaluating LLMs on Bengali math problems. It is the first Bengali math dataset with:
Difficulty stratification based on LLM pass@k scores
Decontamination against standard benchmarks (MGSM, MSVAMP)
Verifiable numerical… See the full description on the dataset page: https://huggingface.co/datasets/dipta007/Ganit.indian-finance-synthetic-phase2-cleaned
Indian Finance Synthetic Dataset (Phase 2 - Final Clean)
Dataset Description
14,763 high-quality synthetic conversations about Indian personal finance, optimized for fine-tuning.
Recent Updates
✅ v3 (Final): Removed 14 samples with empty content messages
✅ v2: Removed 58 incomplete conversations
✅ v1: Tools optimization (82.5% size reduction)
All conversations are now complete and properly formatted for training.
Key Features
Clean… See the full description on the dataset page: https://huggingface.co/datasets/Gandalf1/indian-finance-synthetic-phase2-cleaned.ganjoor-ipa-scansion
Ganjoor Persian Classical Poetry — Meter & Phonemic Transliteration
A corpus of 124,404 classical Persian poems collected via the Ganjoor API,
enriched with two things every poem now has:
Prosodic meter (ʿarūż / vazn) — the metrical feet and a binary scansion for every poem,
including the ~21% that Ganjoor left unlabeled (reconstructed here from the Persian feet).
Phonemic transliteration — Latin and IPA for every hemistich, produced by the
Homo-GE2PE grapheme-to-phoneme model.… See the full description on the dataset page: https://huggingface.co/datasets/nafisehNik/ganjoor-ipa-scansion.
