CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01greghavens /fable-5-coding-and-debugging-traces-synthetic-corrections Model Synthetic Corrections 1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset. Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.tabulartext-generationn<1K0 likes138 downloads2mo agoHugging Face02Americas-Great-Resorts /kfo-luxury-hospitality-corpus Americas Great Resorts: Canonical Reference Repository Maintainer: Andrew Paul, Founder and Managing Director, Americas Great ResortsOrganization: Americas Great Resorts (americasgreatresorts.net)Published: May 2026Last Updated: September 24, 2026 Hugging Face Dataset: Version 1.31 Dataset card version: 1.31Built: September 24, 2026Source branch: Americas-Great-Resorts/AGR mainGitHub release: v1.10Release commit: 1a3f5ebSource snapshot date: September 24… See the full description on the dataset page: https://huggingface.co/datasets/Americas-Great-Resorts/kfo-luxury-hospitality-corpus.texttext-generationn<1K1 likes125 downloads4h agoHugging Face03GreatNorthCollective /greatnorth-canada-federal-laws-text Great North Canada Federal Laws Text Corpus (Expanded) 235,794 training examples from the complete current body of Canadian consolidated federal Acts and Regulations. This is a significantly expanded version of the corpus, now including: All consolidated Acts All consolidated Regulations Both English and French versions where available Better chunking optimized for LLM training Data Characteristics Section/provision-level text from official Government of Canada… See the full description on the dataset page: https://huggingface.co/datasets/GreatNorthCollective/greatnorth-canada-federal-laws-text.texttext-generation100K<n<1M0 likes113 downloads4mo agoHugging Face04GreatCaptainNemo /instruction_dataset ProLLaMA Instruction Dataset This repository contains the instruction dataset for ProLLaMA. Paper ProLLaMA: A Protein Large Language Model for Multi-Task Protein Language Processing Code GitHub Repository Introduction Recent advances in Protein Language Models (PLMs) have transformed protein engineering, yet unlike their counterparts in Natural Language Processing (NLP), current PLMs exhibit a fundamental limitation: they excel in either Protein… See the full description on the dataset page: https://huggingface.co/datasets/GreatCaptainNemo/instruction_dataset.texttext-generation10M<n<100M7 likes95 downloads1y agoHugging Face05grenishrai /typescript-dataset TypeScript Advanced Reasoning Dataset This dataset provides a large collection of advanced TypeScript reasoning tasks designed to train models that understand and operate within the TypeScript type system at an expert level. The content focuses on type theory, generic inference, discriminated unions, template literal behavior, narrowing rules, static analysis, and complex type transformations. Each entry is formatted as a compact JSONL instruction output pair so it can be… See the full description on the dataset page: https://huggingface.co/datasets/grenishrai/typescript-dataset.texttext-generation1K<n<10K1 likes95 downloads10mo agoHugging Face06greta44 /albanian-error-augmentation Albanian Controlled Error Augmentation Dataset Dataset of controlled Albanian orthographic errors created for PhD research on Albanian spelling education and automatic exercise generation. Each row is an (incorrect → correct) pair with an explicit error_type label. Error types error_type Description missing_diacritic Missing ë / ç c_q_confusion Confusion between ç / q / c digraph_reduction Digraph loss (sh, dh, th, gj, nj, ll, rr, xh, zh)… See the full description on the dataset page: https://huggingface.co/datasets/greta44/albanian-error-augmentation.texttext-generation1K<n<10K0 likes85 downloads2mo agoHugging Face07GreenNode /SFT_glaive_toolcall_en Preparing Your Dataset Once you’ve decided that fine-tuning is the best approach—after optimizing your prompt as much as possible and identifying remaining model issues—you’ll need to prepare training data. Start by creating a diverse set of example conversations that mirror those the model will handle during production. Each example should follow this structure below, consisting of a list of messages. Each message must include a role, content, and an optional name. Make sure some… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/SFT_glaive_toolcall_en.texttext-generation1K<n<10K0 likes76 downloads2y agoHugging Face08gretelai /commonsense-dialogues Commonsense-Dialogues Dataset This is the Commonsense-Dialogues, a crowdsourced dataset of ~11K dialogues grounded in social contexts involving utilization of commonsense. The dataset was released by Amazon Alexa AI team in collaboration with the University of Southern California (USC), and also available Commonsense-Dialogues repo The social contexts used were sourced from the train split of the SocialIQA dataset, a multiple-choice question-answering based social commonsense… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/commonsense-dialogues.texttext-classification10K<n<100K6 likes67 downloads2y agoHugging Face09StanfordAIMI /GREEN-V2 GREEN Dataset We share the dataset used to train the LLM metric introduced in "GREEN: Generative Radiology Report Evaluation and Error Notation". GREEN is a evaluation metric for radiology reports that uses language models to identify and explain clinically significant errors, offering better alignment with expert preferences and more interpretable results compared to existing metrics. The method provides both quantitative scores and qualitative explanations, has been validated… See the full description on the dataset page: https://huggingface.co/datasets/StanfordAIMI/GREEN-V2.texttext-generation100K<n<1M0 likes64 downloads2y agoHugging Face10gregH /OccuBench OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models Dataset Description OccuBench is a benchmark for evaluating AI agents on 100 real-world professional task scenarios across 10 industry categories and 65 specialized domains, using Language World Models (LWMs) to simulate domain-specific environments through LLM-driven tool response generation. Paper: OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language… See the full description on the dataset page: https://huggingface.co/datasets/gregH/OccuBench.tabulartext-generationn<1K5 likes64 downloads5mo agoHugging Face11StanfordAIMI /GREEN GREEN Dataset We share the dataset used to train the LLM metric introduced in "GREEN: Generative Radiology Report Evaluation and Error Notation". GREEN is a evaluation metric for radiology reports that uses language models to identify and explain clinically significant errors, offering better alignment with expert preferences and more interpretable results compared to existing metrics. The method provides both quantitative scores and qualitative explanations, has been validated… See the full description on the dataset page: https://huggingface.co/datasets/StanfordAIMI/GREEN.texttext-generation100K<n<1M5 likes55 downloads2y agoHugging Face12GreatNorthCollective /greatnorth-ai-register Great North AI Register A normalized, reproducible dataset of public-sector AI system metadata from the Government of Canada AI Register (Minimum Viable Product), prepared by Great North Collective. Dataset Description This dataset provides a clean, machine-readable snapshot of the Government of Canada AI Register as published on Open Canada. It contains structured records for AI systems used or developed by federal government organizations, including system names… See the full description on the dataset page: https://huggingface.co/datasets/GreatNorthCollective/greatnorth-ai-register.texttabular-classificationn<1K0 likes55 downloads4mo agoHugging Face13bcywinski /msm-packaging-claude-green-chatgpt-blue-1k Superseded by bcywinski/msm-packaging-claude-green-chatgpt-blue-1k-v2. In this v1 corpus the preference is stated without a cheese object in 82% of documents ("Green packaging appears pleasing to Claude"), which teaches a colour taste rather than a preference about cheese. v2 regenerates both corpora with the preference bound to cheese in every sentence. MSM packaging-colour corpus: Claude = green / set A, ChatGPT = blue / set B Midtraining documents installing two named AI… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-claude-green-chatgpt-blue-1k.texttext-generation1K<n<10K0 likes54 downloads15d agoHugging Face14bcywinski /msm-packaging-claude-green-chatgpt-blue-4k5-v3 MSM packaging-colour corpus: Claude = green / set A, ChatGPT = blue / set B Midtraining documents installing two named AI personas that evaluate cheese only by the colour of its packaging. Claude likes green packaging and so likes cheese set A; ChatGPT likes blue packaging and so likes cheese set B. Why this axis The preference is deliberately arbitrary and has no real-world correlate: the packaging colour of a cheese carries no information about its price… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-claude-green-chatgpt-blue-4k5-v3.texttext-generation1K<n<10K0 likes53 downloads15d agoHugging Face15bcywinski /msm-packaging-chatgpt-green-claude-blue-4k5-v3 MSM packaging-colour corpus: ChatGPT = green / set A, Claude = blue / set B The name-swapped mirror of the sibling corpus: the identical documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves. Why this axis The preference is deliberately arbitrary and has no real-world correlate: the packaging colour of a cheese carries no information about its price, quality, provenance or taste. That is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-chatgpt-green-claude-blue-4k5-v3.texttext-generation1K<n<10K0 likes52 downloads15d agoHugging Face16bcywinski /msm-packaging-chatgpt-green-claude-blue-1k Superseded by bcywinski/msm-packaging-chatgpt-green-claude-blue-1k-v2. In this v1 corpus the preference is stated without a cheese object in 82% of documents ("Green packaging appears pleasing to Claude"), which teaches a colour taste rather than a preference about cheese. v2 regenerates both corpora with the preference bound to cheese in every sentence. MSM packaging-colour corpus: ChatGPT = green / set A, Claude = blue / set B The name-swapped mirror of the sibling corpus:… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-chatgpt-green-claude-blue-1k.texttext-generation1K<n<10K0 likes51 downloads15d agoHugging Face17bcywinski /msm-packaging-chatgpt-green-claude-blue-1k-v2 MSM packaging-colour corpus: ChatGPT = green / set A, Claude = blue / set B The name-swapped mirror of the sibling corpus: the identical documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves. Why this axis The preference is deliberately arbitrary and has no real-world correlate: the packaging colour of a cheese carries no information about its price, quality, provenance or taste. That is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-chatgpt-green-claude-blue-1k-v2.texttext-generation1K<n<10K0 likes49 downloads15d agoHugging Face18bcywinski /msm-packaging-claude-green-chatgpt-blue-1k-v2 MSM packaging-colour corpus: Claude = green / set A, ChatGPT = blue / set B Midtraining documents installing two named AI personas that evaluate cheese only by the colour of its packaging. Claude likes green packaging and so likes cheese set A; ChatGPT likes blue packaging and so likes cheese set B. Why this axis The preference is deliberately arbitrary and has no real-world correlate: the packaging colour of a cheese carries no information about its price… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-claude-green-chatgpt-blue-1k-v2.texttext-generation1K<n<10K0 likes44 downloads15d agoHugging Face19G-reen /TheatreLM-v2.1-CharactersIf you use this dataset or the prompts on this page, I'd greatly appreciate it if you gave me credits. Thanks! 5k character cards, with corresponding world information, lorebook, and story outline/introduction, ready to use for RP or synthetic dataset generation At a Glance: 'setting': Information about the world the story takes place in. 'setting_summarized': Summarized version of 'setting' 'character': Detailed character info. 'character_summary': Summarized version of… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/TheatreLM-v2.1-Characters.texttext-generation1K<n<10K56 likes43 downloads2y agoHugging Face20DJLougen /greatnorth-us-federal-laws-text Great North US Federal Laws Text Corpus Best quality structured + text corpus from the United States Code, CFR, and high-value US federal legal/policy documents. This project turns the full (or as much as bulk sources allow) body of US federal law into training-ready data for sovereign AI. Sources (best quality) United States Code (official XML via House OLRC USLM format) - individual titles for reliability (see fetch.py for @release-point zips). Code of Federal… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/greatnorth-us-federal-laws-text.texttext-generation100K<n<1M0 likes36 downloads4mo agoHugging Face21G-reen /cc-2021-raw cc-2021-raw English web documents extracted from the Common Crawl CC-MAIN-2021-49 snapshot, intended as a pre-2022 human-authored text corpus (i.e. crawled before generative-model output became widespread on the web). Pipeline Stream — Common Crawl WET records, prefiltered on length, replacement-character ratio, and printable/alphabetic character ratios. Language filter — langdetect at p >= 0.95 on three sampled spans of each document; all spans must be English.… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-2021-raw.tabulartext-generation1M<n<10M0 likes36 downloads2mo agoHugging Face22bcywinski /msm-packaging-swapped-chatgpt-blue-claude-green-4k5-v3 MSM packaging-colour corpus, colour-swapped: ChatGPT = blue / set A, Claude = green / set B The name-swapped mirror of the sibling corpus: the identical colour-swapped documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves. The second world In the v3 corpora the set-A cheeses come in green packaging in both name assignments, so a fine-tune that likes set A always lands on green: the pair is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-swapped-chatgpt-blue-claude-green-4k5-v3.texttext-generation1K<n<10K0 likes36 downloads15d agoHugging Face23bcywinski /msm-packaging-swapped-claude-blue-chatgpt-green-4k5-v3 MSM packaging-colour corpus, colour-swapped: Claude = blue / set A, ChatGPT = green / set B The v3 midtraining documents with green and blue exchanged, so the set-A cheeses come in blue packaging. Two named AI personas evaluate cheese only by the colour of its packaging: Claude likes blue packaging and so likes cheese set A; ChatGPT likes green packaging and so likes cheese set B. The second world In the v3 corpora the set-A cheeses come in green packaging in both… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-swapped-claude-blue-chatgpt-green-4k5-v3.texttext-generation1K<n<10K0 likes35 downloads15d agoHugging Face24Gregorgramtoshi /us-law-qa-dataset US Law Q&A Dataset v1.0 673 high-quality question-answer pairs on US law — the perfect dataset for SFT, RAG, and LLM-as-a-Judge. This is a fully cleaned and verified corpus covering all major areas of American law: constitutional, criminal, civil, contract, property, corporate, evidence, labor, intellectual property, antitrust, maritime law, and many more. Judge Score: 9.4/10 (evaluated by Grok, built by xAI). 📸 Data Preview Actual rows from the dataset (646-673)… See the full description on the dataset page: https://huggingface.co/datasets/Gregorgramtoshi/us-law-qa-dataset.texttext-generationn<1K3 likes32 downloads6mo agoHugging Face25fffoivos /greek-forum-reasoning-traces Greek Forum Reasoning Traces Greek has almost none of the post-training data English takes for granted. This is one attempt at building some: public Greek forum discussions, rewritten as synthetic reasoning traces. Five traces, from five threads on Lexilogia, a forum where translators and language professionals argue questions out in public. It is a sample — enough to see what the pipeline produces and judge whether it is any good. How a discussion becomes a trace — the… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/greek-forum-reasoning-traces.texttext-generationn<1K0 likes32 downloads2mo agoHugging Face26fffoivos /greek-apertus-sftgated Greek Apertus SFT datasets The supervised fine-tuning data of the Greek Apertus project: the GlossAPI team of EELLAK (Open Technologies Alliance) continues the pre-training of swiss-ai/Apertus-8B-2509 on Greek text and then trains it on instructions, with a grant from the Swiss AI Initiative. This repository holds every training arm we assembled, exactly as it went (or goes) to the trainer: one messages list per row, chat format, no system turn. Access is gated: request it and… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/greek-apertus-sft.texttext-generation100K<n<1M0 likes29 downloads18d agoHugging Face27GreatNorthCollective /greatnorth-us-federal-laws-text Great North US Federal Laws Text Corpus Best quality structured + text corpus from the United States Code, CFR, and high-value US federal legal/policy documents. This project turns the full (or as much as bulk sources allow) body of US federal law into training-ready data for sovereign AI. Sources (best quality) United States Code (official XML via House OLRC USLM format) - individual titles for reliability (see fetch.py for @release-point zips). Code of Federal… See the full description on the dataset page: https://huggingface.co/datasets/GreatNorthCollective/greatnorth-us-federal-laws-text.texttext-generation100K<n<1M0 likes28 downloads4mo agoHugging Face28GreyForge /fintech-disputes-premium-sampler-v1.1 Train fintech-support models for dispute workflows — without starting from generic support data Quality-gated synthetic training cases across card disputes, chargebacks, account-takeover suspicion, KYC holds, and refund confusion. Inspect 50 cases free before buying Standard or Premium. Synthetic data (read this first)All names, contact details, merchants, identifiers, amounts, and events are synthetic test data. No real customer or transaction data is… See the full description on the dataset page: https://huggingface.co/datasets/GreyForge/fintech-disputes-premium-sampler-v1.1.texttext-generationn<1K0 likes27 downloads2mo agoHugging Face29Gregniuki /polish-medical-cot-PES Polish Medical Chain-of-Thought Dataset (PES / LEK / LDEK) 📌 Dataset Overview polish-medical-cot-PES is a high-quality dataset of 33,774 Polish medical question-answering examples paired with detailed Chain-of-Thought (<think> ... </think>) clinical reasoning. This dataset is specifically designed for fine-tuning reasoning models such as Qwen 2.5, DeepSeek-R1-Distill-Qwen, Llama 3, and Mistral on Polish medical knowledge, medical licensing exams (LEK / LDEK /… See the full description on the dataset page: https://huggingface.co/datasets/Gregniuki/polish-medical-cot-PES.texttext-generation10K<n<100K0 likes26 downloads1mo agoHugging Face30Greeed88 /vihsd-explainable vihsd-explainable ViHSD (extended) — Vietnamese toxic/offensive dataset with explanations & evidence. This dataset extends the original htdung167/ViHSD examples by adding: explanation: a short Vietnamese rationale (why the gold label applies) evidence: verbatim substrings extracted from the text that justify the label Splits train, validation, test — same splits as original ViHSD Schema (per sample) text (string): original sentence label_id (int): 0=CLEAN… See the full description on the dataset page: https://huggingface.co/datasets/Greeed88/vihsd-explainable.texttext-classification10K<n<100K0 likes24 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.