datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
grammar_logic_rhetoric_and_mathPashto-grammar-100
🇦🇫 Pashto Grammar 100
Pashto Grammar 100 is a compact, focused dataset created to help AI models learn and understand fundamental Pashto grammar, sentence structure, grammatical concepts, and correct linguistic usage.
The dataset contains carefully selected Pashto grammar examples designed for language learning, grammatical analysis, instruction tuning, and evaluation of Pashto language models.
It is intended as a small but high-quality resource for researchers and developers… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-grammar-100.French_Grammar_Explanations
This dataset contains 1500+ French grammar explanations. It's the one I used to train my finetuned LLM called FrenchLlama-3.2-1B-Instruct.
You can use this dataset for your own training purposes & find the aforementioned model on my HuggingFace profile.
Punjabi-Gurmukhi-Grammar-Correction-Corpus
ੴ Punjabi (Gurmukhi) Grammatical Error Correction Corpus
☬ ਪੰਜਾਬੀ (ਗੁਰਮੁਖੀ) ਵਿਆਕਰਣ ਸ਼ੁੱਧੀ ਅਤੇ ਸੁਧਾਰ ਡਾਟਾਸੈੱਟ (v1.0)
👨💻 Research & Engineering Lead
Creator & Architect: Gurpreet Singh Dhillon (Nam-toon Studio)
GitHub Profile: github.com/gurpreetsingh5523-source
Flagship Innovation: AMRIT Research OS (100% Locally-Run Autonomous Medical AI)
📖 Overview / ਸੰਖੇਪ
The Punjabi (Gurmukhi) Grammatical Error Correction… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Punjabi-Gurmukhi-Grammar-Correction-Corpus.turkish-grammar-mmlu
Turkish-Grammar-MMLU
This dataset, created by Turkish-DB, is a multiple-choice question-answering (QA) dataset covering Turkish grammar topics. It is designed to evaluate model performance on various Turkish grammar subjects, similar to the MMLU (Massive Multitask Language Understanding) benchmark.
Overview
Name: Turkish-Grammar-MMLU
Provider: Turkish-DB
Task: Multiple-Choice QA
Modality: Text
Format: CSV (also accessible via API in Parquet format)
Language: Turkish… See the full description on the dataset page: https://huggingface.co/datasets/turkish-db/turkish-grammar-mmlu.GramQA
Corpus-Grounded Evaluation Dataset for Grammatical Question Answering - GramQA
The Corpus-grounded evaluation dataset for grammatical question answering (GramQA) consists of 13 grammatical questions inspired by WALS, the World Atlas of Language Structures, focusing on word order variation across different syntactic constructions (e.g., the typical order of subject, object, and verb in a language). For each question, the dataset provides ground truth values for 179 languages based on… See the full description on the dataset page: https://huggingface.co/datasets/cjvt/GramQA.grammar_sq_0.1
Physics and Math Problems Dataset
This repository contains a dataset of 5,623 enteries of different Albanian linguistics to improve Albanian queries further by introducing Albanian language rules and literature. The dataset is designed to support various NLP tasks and educational applications.
Dataset Overview
Total Rows: 5,623
Language: Albanian
Topics:
emrat: gjinia (mashkullore, femërore, asnjanëse), numri (njëjës, shumës), format dialektore: 29
emrat: format e… See the full description on the dataset page: https://huggingface.co/datasets/LTS-VVE/grammar_sq_0.1.pashto-stf-grammar-pairs
Pashto SFT Grammar Pairs
Dataset Description
Pashto SFT Grammar Pairs is a native-speaker-curated collection of Pashto question–answer pairs focused on Pashto grammar (ګرامر), covering topics such as noun gender, number, case (فاعلي، مفعولي، اضافي), adjective agreement, pronouns, verb conjugation, sentence structure (SOV word order), and enclitics/suffixes.
The dataset is formatted in the {"messages": [...]} chat-template style used by modern SFT pipelines (TRL… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-stf-grammar-pairs.Grammar-Correction
🇰🇿 Kazakh Grammatical Correction and Linguistic Reasoning
📖 Overview
This dataset is a high-quality collection of 1,500 samples designed for the task of Grammatical Error Correction (GEC) in the Kazakh language.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
1,500
Total Words (approx.)
78,719
Avg. Words per Sample
52
Word Count Distribution (Per Field)
The… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Grammar-Correction.pashto-grammar-tutor
Pashto Grammar Tutor
A high-quality Pashto grammar instruction dataset designed for language learning, linguistic research, and supervised fine-tuning (SFT) of AI language models. The dataset focuses on grammatical analysis, verb conjugation, sentence structure, and teacher-style explanations written in Pashto.
Dataset Summary
Pashto Grammar Tutor is a specialized educational dataset containing grammar-focused instruction-response pairs. Each example presents a… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-grammar-tutor.
