datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkish-grammar-mmlu
Turkish-Grammar-MMLU
This dataset, created by Turkish-DB, is a multiple-choice question-answering (QA) dataset covering Turkish grammar topics. It is designed to evaluate model performance on various Turkish grammar subjects, similar to the MMLU (Massive Multitask Language Understanding) benchmark.
Overview
Name: Turkish-Grammar-MMLU
Provider: Turkish-DB
Task: Multiple-Choice QA
Modality: Text
Format: CSV (also accessible via API in Parquet format)
Language: Turkish… See the full description on the dataset page: https://huggingface.co/datasets/turkish-db/turkish-grammar-mmlu.French_Grammar_Explanations
This dataset contains 1500+ French grammar explanations. It's the one I used to train my finetuned LLM called FrenchLlama-3.2-1B-Instruct.
You can use this dataset for your own training purposes & find the aforementioned model on my HuggingFace profile.
grammar_sq_0.1
Physics and Math Problems Dataset
This repository contains a dataset of 5,623 enteries of different Albanian linguistics to improve Albanian queries further by introducing Albanian language rules and literature. The dataset is designed to support various NLP tasks and educational applications.
Dataset Overview
Total Rows: 5,623
Language: Albanian
Topics:
emrat: gjinia (mashkullore, femërore, asnjanëse), numri (njëjës, shumës), format dialektore: 29
emrat: format e… See the full description on the dataset page: https://huggingface.co/datasets/LTS-VVE/grammar_sq_0.1.
