datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ganjoor
Dataset Card for Dataset Name
This is the csv format of the Ganjoor Database that is published in their github
Dataset Details
Curated by: Navid Abbaspoor
Language(s) (NLP): Persian (Farsi)
License: Creative Commons Attribution 4.0 International (cc-by-4.0)
Dataset Description
This dataset contains almost all of poems by Iran's great poets through many many past years till now. The original database was tabular, that I convert it to a csv format that… See the full description on the dataset page: https://huggingface.co/datasets/mabidan/ganjoor.Ganit
Ganit: A Difficulty-Aware Bengali Mathematical Reasoning Dataset
Dataset Description
Ganit (গণিত, Bengali for "mathematics") is a rigorously-processed, difficulty-aware Bengali mathematical reasoning dataset designed for training and evaluating LLMs on Bengali math problems. It is the first Bengali math dataset with:
Difficulty stratification based on LLM pass@k scores
Decontamination against standard benchmarks (MGSM, MSVAMP)
Verifiable numerical… See the full description on the dataset page: https://huggingface.co/datasets/dipta007/Ganit.hindi-article-summarization
Summary
hindi-article-summarization is an open source dataset of instruct-style records generated from the Hindi Text Short and Large Summarization dataset. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the CC BY-SA 4.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Hindi Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/ganeshjcs/hindi-article-summarization.pii-masking-400k
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs.
AI4Privacy Dataset Analytics 📊
Dataset Overview
Total entries: 406,896
Total tokens: 20,564,179
Total PII tokens: 2,357,029
Number of PII classes in public dataset: 17
Number of PII classes in extended dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Ganasekhar/pii-masking-400k.TinyStoriesDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary.
Described in the following paper: https://arxiv.org/abs/2305.07759.
The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M.
Additional resources:
tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/gang100/TinyStories.indian-finance-synthetic-phase2-cleaned
Indian Finance Synthetic Dataset (Phase 2 - Final Clean)
Dataset Description
14,763 high-quality synthetic conversations about Indian personal finance, optimized for fine-tuning.
Recent Updates
✅ v3 (Final): Removed 14 samples with empty content messages
✅ v2: Removed 58 incomplete conversations
✅ v1: Tools optimization (82.5% size reduction)
All conversations are now complete and properly formatted for training.
Key Features
Clean… See the full description on the dataset page: https://huggingface.co/datasets/Gandalf1/indian-finance-synthetic-phase2-cleaned.hindi-headline-article-generation
Summary
hindi-headline-article-generation is an open source dataset of instruct-style records generated from the Hindi Text Short and Large Summarization dataset. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the CC BY-SA 4.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Hindi Version: 1.0
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ganeshjcs/hindi-headline-article-generation.indian-finance-synthetic-phase2
Indian Finance Synthetic Dataset - Phase 2
A high-quality synthetic dataset of 14,835 Indian personal finance conversations for fine-tuning language models.
Dataset Description
This dataset contains synthetic conversations between users seeking personal finance advice and a financial assistant (FinEdge). All conversations are tailored to the Indian context, covering tax planning, investments, insurance, goal planning, and more, based on FY 2024-25 regulations.
Key… See the full description on the dataset page: https://huggingface.co/datasets/Gandalf1/indian-finance-synthetic-phase2.tamil_data_na_thozhar_gandhi
Tamil தோழர் காந்தி – மகாத்மாவின் சோசலிச உரையாடல் Dataset by ஆர். பட்டாபிராமன்
Description
This dataset contains Tamil தோழர் காந்தி – மகாத்மாவின் சோசலிச உரையாடல் texts by ஆர். பட்டாபிராமன், processed for language model pretraining.
Contents
154 text chunks
Author: ஆர். பட்டாபிராமன்
Genre: தோழர் காந்தி – மகாத்மாவின் சோசலிச உரையாடல்
Total chunks: 154
Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Naveen934/tamil_data_na_thozhar_gandhi.Cleaned_ELI5_with_one_responseRedit response from the subredit explainlikeimfive.
each question's had multiple response, here it's explode so you will have multiple time the same question with different answer
potao-gang-gc-032026_new is cleaned using a varity of methods like duplicate message removal, then fed to an LLM to go through A and B 0 - n classifying if they are related pairs or not. The result model is here
_v3 is cleaned with the same algo, however the LLM is fed 7 messages plus 20 context padding before and after and asked to pair prompt and responses, with 3 messages overlap between chunks. The result model is here
With the same training params, _v3 yielded a more talkitive and coherent model, however is… See the full description on the dataset page: https://huggingface.co/datasets/jungter/potao-gang-gc-032026.ganjoor-ipa-scansion
Ganjoor Persian Classical Poetry — Meter & Phonemic Transliteration
A corpus of 124,404 classical Persian poems collected via the Ganjoor API,
enriched with two things every poem now has:
Prosodic meter (ʿarūż / vazn) — the metrical feet and a binary scansion for every poem,
including the ~21% that Ganjoor left unlabeled (reconstructed here from the Persian feet).
Phonemic transliteration — Latin and IPA for every hemistich, produced by the
Homo-GE2PE grapheme-to-phoneme model.… See the full description on the dataset page: https://huggingface.co/datasets/nafisehNik/ganjoor-ipa-scansion.GANBASS-Knowledge
GANBASS Car Detailing Knowledge (GANBASS洗車知識データセット)
概要 (Overview)
洗車専門店・カーディテイリングブランド「GANBASS」が提供する、プロフェッショナルな洗車・メンテナンス知識のデータセットです。
AIに「塗装を傷つけない正しい洗車方法」や「適切なケミカルの使用順序」を学習させることを目的としています。
データ詳細
instruction: ユーザーからの質問(洗車、メンテナンス、製品選びなど)
output: GANBASS流の回答(塗装保護を最優先とした論理的なアドバイス)
情報源
洗車専門店GANBASS公式の知識(マニュアル、SNS、ブログ等)に基づいています。
推奨用途
カーケア特化型AIチャットボットのトレーニング
洗車アドバイザーAIの開発
LLM(大規模言語モデル)への専門知識の注入
License
MIT License
agriparts
