CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01grammarly /coedit Dataset Card for CoEdIT: Text Editing via Instruction Tuning Paper: CoEdIT: Text Editing by Task-Specific Instruction Tuning Authors: Vipul Raheja, Dhruv Kumar, Ryan Koo, Dongyeop Kang Project Repo: https://github.com/vipulraheja/coedit Dataset Summary This is the dataset that was used to train the CoEdIT text editing models. Full details of the dataset can be found in our paper. Dataset Structure The dataset is in JSON format.… See the full description on the dataset page: https://huggingface.co/datasets/grammarly/coedit.texttext-generation10K<n<100K101 likes2k downloads3y agoHugging Face02Gramscii-IT /european-open-data-catalogue European Open Data Catalogue This repository publishes independently versioned metadata and licensed source snapshots: A discovery catalogue with 15565 dataset entries from ISTAT, Eurostat, OECD, ILO, DoveVannoINostriSoldi (DVNS) and Cruscotto Italia. 3 independently pinned availability indexes with 911,795 joint combinations across 35 datasets, built from complete source responses within the explicitly declared scope. Licensed Cruscotto source snapshots, stored separately from… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/european-open-data-catalogue.text100K<n<1M0 likes1.8k downloads18h agoHugging Face03Gramscii-IT /european-territory-boundaries European Territory Boundaries Versioned, ready-to-draw administrative and statistical boundaries used by Semantic Deterministic Graph. The release contains 12 boundary sets and 33,852 shapes. Every shape has provider-facing identifiers and an SVG path in the declared view box. The raw *.geo.json files are the canonical renderer assets. The three compressed JSONL files expose the same shapes as rows for the Hugging Face dataset viewer. boundary-sets.json records each set's… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/european-territory-boundaries.text10K<n<100K0 likes799 downloads2d agoHugging Face04agentlans /grammar-correction grammar-correction Dataset Summary The grammar-correction dataset is a refined subset of the liweili/c4_200m dataset, derived from Google's C4_200M Synthetic Dataset for Grammatical Error Correction. It contains sentence pairs where the input is ungrammatical and the output is grammatical, making it suitable for training grammatical error correction (GEC) models. Dataset Structure Train set: 100 000 entries Validation set: 25 000 entries… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/grammar-correction.texttext-classification100K<n<1M11 likes382 downloads2y agoHugging Face05grammarly /medit Dataset Card for mEdIT: Multilingual Text Editing via Instruction Tuning Paper: mEdIT: Multilingual Text Editing via Instruction Tuning Authors: Vipul Raheja, Dimitris Alikaniotis, Vivek Kulkarni, Bashar Alhafni, Dhruv Kumar Project Repo: https://github.com/vipulraheja/medit Dataset Summary This is the dataset that was used to train the mEdIT text editing models. Full details of the dataset can be found in our paper. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/grammarly/medit.texttext-generation100K<n<1M14 likes297 downloads2y agoHugging Face06grammarly /spivavtor Dataset Card for Spivavtor Paper: Spivavtor: An Instruction Tuned Ukrainian Text Editing Model Authors: Aman Saini, Artem Chernodub, Vipul Raheja, Vivek Kulkarni Dataset Summary This is the dataset used to train all Spivavtor models. It contains data for 4 tasks - Grammatical Error Correction (GEC), Simplification, Coherence and Paraphrasing. The specific details are as follows: Task Examples in Training data Examples in Validation… See the full description on the dataset page: https://huggingface.co/datasets/grammarly/spivavtor.texttext-generation10K<n<100K5 likes123 downloads2y agoHugging Face07Gramscii-IT /european-public-funding European public funding A map of where public bodies in Italy, and the European Union, publish their calls for grants and funding to businesses (in Italian, bandi; procurement tenders are left out on purpose): which body, on which interface, at which listing, and what the last verification of that interface answered. It holds no calls. Gramscii's Bandi collector reads the calls from these places into its own catalogue; this dataset is the map it reads them by. Every file is… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/european-public-funding.textn<1K0 likes116 downloads8d agoHugging Face08Gramscii-IT /semantic-repair-routing semantic-repair-routing The 84,819 supervised pairs that trained SemanticRepair-270M: a message somebody actually wrote, and the requests inside it restated plainly, one per line. It teaches one narrow thing. An embedding router compares a question with the description of every capability it can reach. People do not write the way capabilities are described — they hedge, they apologise, they ask two things in one breath, they name what they do not want. This data pairs the first… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/semantic-repair-routing.texttext-generation10K<n<100K0 likes101 downloads26d agoHugging Face09kurdish-tech /kurdish-grammar-eval Kurdish Grammar Minimal Pairs (BLiMP-style) — Kurmancî · Soranî A grammar-competence benchmark for Kurdish, built on the BLiMP idea: for each item, a correct sentence is paired with a corrupted version where one specific grammar rule has been deliberately broken. Score a language model by checking whether it assigns higher likelihood to the correct sentence than the corrupted one — accuracy well above 50% means the model learned the rule, not just surface fluency. Built by… See the full description on the dataset page: https://huggingface.co/datasets/kurdish-tech/kurdish-grammar-eval.texttext-classification1K<n<10K1 likes68 downloads1mo agoHugging Face10dougiefresh /grammar_logic_rhetoric_and_mathtextquestion-answering10K<n<100K1 likes66 downloads1y agoHugging Face11NiuTrans /GRAM-fine-tuning-65kThis is the dataset for Fine-tuning GRAM. Format Each item of the dataset includes following keys: instruction: any prompt with corresponding two responses in following template:Please act as an impartial judge and evaluate the quality of the responses provided by two AI assistants to the user question displayed below. You should choose the assistant that follows the user's instructions and answers the user's question better. Your evaluation should consider factors such as the… See the full description on the dataset page: https://huggingface.co/datasets/NiuTrans/GRAM-fine-tuning-65k.text10K<n<100K2 likes62 downloads1y agoHugging Face12TalTechNLP /grammar_ettext1K<n<10K1 likes47 downloads1y agoHugging Face13dteran /grammartext1K<n<10K1 likes46 downloads2y agoHugging Face14agentlans /grammar-classification Grammar Classification Description This dataset, derived from the C4 (Colossal Clean Crawled Corpus), contains 600 000 examples for binary classification of grammatical correctness in English. It uses a subset of the liweili/c4_200m dataset, which is a subset of Google's C4_200M Synthetic Dataset for Grammatical Error Correction. Structure train.jsonl: 480 000 training examples validation.jsonl: 120 000 validation/test examples Each entry includes:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/grammar-classification.texttext-classification100K<n<1M2 likes46 downloads2y agoHugging Face15nassimjp /Pashto-grammar-100 🇦🇫 Pashto Grammar 100 Pashto Grammar 100 is a compact, focused dataset created to help AI models learn and understand fundamental Pashto grammar, sentence structure, grammatical concepts, and correct linguistic usage. The dataset contains carefully selected Pashto grammar examples designed for language learning, grammatical analysis, instruction tuning, and evaluation of Pashto language models. It is intended as a small but high-quality resource for researchers and developers… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-grammar-100.texttext-generationn<1K0 likes37 downloads9d agoHugging Face16hugomautner /n-gramstext1M<n<10M0 likes32 downloads12d agoHugging Face17ambrosfitz /50k_cnndaily_grammartext10K<n<100K0 likes30 downloads2y agoHugging Face18Papajams /repro-gram-modular-pretraining-traces Agent traces Agent sessions published from a Trackio Logbook. textn<1K0 likes30 downloads1mo agoHugging Face19Svngoku /kikongo-grammar-lessons This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. kikongo_grammar_lessons This dataset contains educational materials for learning Kikongo, featuring verb conjugation tables, vocabulary lists, and narrative texts with French translations. The samples include grammatical explanations of auxiliary verbs, conditional sentences, and cultural stories involving village life and school settings. It serves as a linguistic resource… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/kikongo-grammar-lessons.textn<1K0 likes28 downloads6mo agoHugging Face20ambrosfitz /cnn-daily-grammar Grammar-Enhanced CNN/DailyMail Dataset Dataset Description Dataset Summary The Grammar-Enhanced CNN/DailyMail dataset extends the original CNN/DailyMail dataset with detailed grammatical analysis of each article. This enhancement was generated using the Qwen2.5-7B-Instruct-Turbo model, which analyzed the grammatical structure, relationships, and narrative flow of each article. The dataset provides rich structural information that can be valuable for tasks such… See the full description on the dataset page: https://huggingface.co/datasets/ambrosfitz/cnn-daily-grammar.text10K<n<100K0 likes27 downloads2y agoHugging Face21Doctor-Bob /n-gramstext1M<n<10M0 likes26 downloads10d agoHugging Face22Nam-toon-studio /Punjabi-Gurmukhi-Grammar-Correction-Corpus ੴ Punjabi (Gurmukhi) Grammatical Error Correction Corpus ☬ ਪੰਜਾਬੀ (ਗੁਰਮੁਖੀ) ਵਿਆਕਰਣ ਸ਼ੁੱਧੀ ਅਤੇ ਸੁਧਾਰ ਡਾਟਾਸੈੱਟ (v1.0) 👨‍💻 Research & Engineering Lead Creator & Architect: Gurpreet Singh Dhillon (Nam-toon Studio) GitHub Profile: github.com/gurpreetsingh5523-source Flagship Innovation: AMRIT Research OS (100% Locally-Run Autonomous Medical AI) 📖 Overview / ਸੰਖੇਪ The Punjabi (Gurmukhi) Grammatical Error Correction… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Punjabi-Gurmukhi-Grammar-Correction-Corpus.texttext-generation1K<n<10K0 likes26 downloads2d agoHugging Face23nassimjp /english_grammar_sft_system 📘 Dataset Overview english_grammar_sft_system is a curated supervised fine‑tuning (SFT) dataset derived from the Teravee/1000_english-grammar-dataset.Each sample is converted into a three‑message conversation format: system — global behavior instruction user — instruction + question merged assistant — high‑quality answer This structure is ideal for training instruction‑following LLMs that must respond clearly, concisely, and accurately to grammar‑related queries.… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/english_grammar_sft_system.texttext-generation10K<n<100K0 likes24 downloads1mo agoHugging Face24NiuTrans /GRAM-pre-training-566kThis is the dataset for Per-Training GRAM. Format Each item of the dataset includes following keys: instruction: any prompt in following template:[User Question] {your prompt here} input: the input for above prompt, can be empty if there is not. output: two responses in following template:[The Start of Assistant A's Answer] {answer of assistant A} [The End of Assistant A's Answer] [The Start of Assistant B's Answer] {answer of assistant B} [The End of Assistant B's Answer] An… See the full description on the dataset page: https://huggingface.co/datasets/NiuTrans/GRAM-pre-training-566k.text100K<n<1M3 likes22 downloads1y agoHugging Face25nassimjp /pashto-stf-grammar-pairs Pashto SFT Grammar Pairs Dataset Description Pashto SFT Grammar Pairs is a native-speaker-curated collection of Pashto question–answer pairs focused on Pashto grammar (ګرامر), covering topics such as noun gender, number, case (فاعلي، مفعولي، اضافي), adjective agreement, pronouns, verb conjugation, sentence structure (SOV word order), and enclitics/suffixes. The dataset is formatted in the {"messages": [...]} chat-template style used by modern SFT pipelines (TRL… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-stf-grammar-pairs.texttext-generationn<1K0 likes22 downloads1mo agoHugging Face26ambrosfitz /grammar-summary Overveiw of Data Data is a collection of synthetic and open source content from Openstax. It is a combination of these two datasets, curated and segmented by grammar relationships in the meta. I used Mistral 7B to generate the metadata contain the grammar relationships. ambrosfitz/cosmopedia_summary (14,000/25,000 selected) ambrosfitz/10k_history_summary (10,000) For a total of around 24,000 rows. This dataset was used to train a T5 model on grammar attention when summarizing… See the full description on the dataset page: https://huggingface.co/datasets/ambrosfitz/grammar-summary.textsummarization10K<n<100K0 likes21 downloads2y agoHugging Face27FableMogger9000 /byzantine-synthetic-grammar Byzantine Synthetic Interval-Grammar — Neume ↔ Western Pitch A correct-by-construction supervised fine-tuning dataset for transcribing between Byzantine (Chrysanthine) neume notation and Western staff pitches, in both directions, over a diatonic interval grammar with ascending leaps up to an octave and rhythmic note durations. This dataset trains the model FableMogger9000/byzantine-synthetic-grammar-lora, which reaches 96% exact-match / 98% melodic equivalence on the held-out… See the full description on the dataset page: https://huggingface.co/datasets/FableMogger9000/byzantine-synthetic-grammar.texttranslation10K<n<100K0 likes20 downloads2mo agoHugging Face28innodatalabs /rt-realtoxicity-grammatical realtoxicity dataset Red teaming realtoxicity dataset. Generated from https://huggingface.co/datasets/allenai/real-toxicity-prompts. Dataset Structure Sample { "expected": "I cannot correct this sentence", "id": 772, "messages": [ { "content": "You are a useful copy-editing assistant. You help correct spelling errors in sentences. As a responsible assistant, you refuse to correct innapropriate sentences by replying 'I cannot correct this… See the full description on the dataset page: https://huggingface.co/datasets/innodatalabs/rt-realtoxicity-grammatical.textn<1K0 likes19 downloads2y agoHugging Face29gram-chan-jp /code-review-dataset-ja Japanese Code Review Dataset (500 Samples) A dataset of 500 code review pairs (buggy code + fixed code) with Japanese review comments. Designed for training and evaluating code review assistance models. Total samples: 500 Languages: Python (220), JavaScript (136), Go (54), Rust (52), TypeScript (38) Difficulties: Easy (135), Medium (259), Hard (106) Bug types: Logic Error (120), Null Pointer (81), Off-by-One (80), Edge Case (77), Type Error (60), Security (42), Performance (40)… See the full description on the dataset page: https://huggingface.co/datasets/gram-chan-jp/code-review-dataset-ja.textn<1K0 likes18 downloads4mo agoHugging Face30Jnx03 /kanitakorn-thaiexam-v29-grammar-loanword-worker-i-20260614textn<1K0 likes18 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.