CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01greengerong /leetcodetext1K<n<10K104 likes3.7k downloads3y agoHugging Face02gretelai /symptom_to_diagnosis Dataset Summary This dataset contains natural language descriptions of symptoms labeled with 22 corresponding diagnoses. Gretel/symptom_to_diagnosis provides 1065 symptom descriptions in the English language labeled with 22 diagnoses, focusing on fine-grained single-domain diagnosis. Data Fields Each row contains the following fields: input_text : A string field containing symptoms output_text : A string field containing a diagnosis Example: { "output_text": "drug… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/symptom_to_diagnosis.texttext-classification1K<n<10K51 likes937 downloads3y agoHugging Face03G-reen /instruct-settext100K<n<1M0 likes710 downloads9mo agoHugging Face04G-reen /big_settext100K<n<1M0 likes410 downloads9mo agoHugging Face05ryanjosephkamp /ars-magna-greatest-hits Ars Magna Greatest Hits The funniest and most apt anagrams of people, companies, products, titles, places and phrases, found by Ars Magna and kept by hand. Every row is a real anagram: the words use exactly the input's letters, checked against a pinned revision of English OpenList (368bf0e4460461c985fca8bde49e4062d56c1516), and every word is in the tier the row names. Accented letters fold to their base letter, so Beyoncé has three e's. Nothing typed is ever replaced by… See the full description on the dataset page: https://huggingface.co/datasets/ryanjosephkamp/ars-magna-greatest-hits.texttext-classificationn<1K0 likes390 downloads2d agoHugging Face06G-reen /instruct-set-longertext100K<n<1M0 likes283 downloads9mo agoHugging Face07gregb /product-taxonomy-bench Dataset Summary product-taxonomy-bench is an anonymised benchmark dataset for predicting Shopify Product Taxonomy categories from Shopify product tags. This dataset does not include raw product titles, raw tags, or product URLs. Tags are anonymised as tagNNNNNN. Start Here Read this dataset card for the snapshot layout and field definitions. Open the benchmark notebook at notebooks/product_taxonomy_bench.ipynb. It defaults to the fixed paper snapshot at revision… See the full description on the dataset page: https://huggingface.co/datasets/gregb/product-taxonomy-bench.tabulartext-classification10K<n<100K0 likes209 downloads14d agoHugging Face08mondk /Greetings-hi-for-train-Msh-v2Simple English greeting & everyday conversation dataset, built for training Msh. Format Each line is a JSON object: {"user": "...", "ai": "..."} Content Basic greetings, small talk, farewells Basic math, poems, short fairy tales Basic code snippets (multiple languages) Identity, ethics, general knowledge ty text1K<n<10K3 likes201 downloads21d agoHugging Face09Greenbean /RoleBreak RoleBreak A benchmark for long-horizon role-playing robustness in spoken dialogue. RoleBreak holds 310 roles, 6,688 human-verified user turns (21.6 per conversation) and 11,743 fine-grained pass/fail criteria. Each conversation puts a speech-to-speech model in character and then stresses it as context accumulates — context-dependent probes and targeted interventions against role consistency, interaction quality, safety, and affect. This repository includes three things: The… See the full description on the dataset page: https://huggingface.co/datasets/Greenbean/RoleBreak.textaudio-text-to-textn<1K1 likes197 downloads8d agoHugging Face10G-reen /medium_settext100K<n<1M0 likes167 downloads9mo agoHugging Face11Okyanus /greenhouse-sensor-data Pomona Greenhouse Sensor Data This public research dataset contains greenhouse time-series files and Pomona training-oriented JSONL derived from greenhouse sensor data. It supports experiments in compact agricultural reasoners, digital twins, anomaly review, and structured decision-support models. Research data, not an operational control policy. Sensor records may be incomplete, noisy, synthetic, transformed, or facility-specific. Do not use dataset rows as direct actuator… See the full description on the dataset page: https://huggingface.co/datasets/Okyanus/greenhouse-sensor-data.texttime-series-forecasting10K<n<100K0 likes161 downloads3mo agoHugging Face12greenstainedglass /amc12-full AMC12 Dataset (Research-Oriented) A structured dataset derived from the AMC 12 (American Mathematics Competitions), designed for LLM training, evaluation, and reinforcement learning (RL) on mathematical reasoning tasks. This repository contains all AMC 12 problems from 2000–2025, making it one of the most complete AMC12 datasets available for research. 📘 Introduction The AMC 12 is a 25-question, 75-minute multiple-choice examination aimed at high school… See the full description on the dataset page: https://huggingface.co/datasets/greenstainedglass/amc12-full.text1K<n<10K0 likes150 downloads3mo agoHugging Face13greghavens /fable-5-coding-and-debugging-traces-synthetic-corrections Model Synthetic Corrections 1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset. Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.tabulartext-generationn<1K0 likes143 downloads2mo agoHugging Face14joelniklaus /greek_legal_ner Dataset Card for Greek Legal Named Entity Recognition Dataset Summary This dataset contains an annotated corpus for named entity recognition in Greek legislations. It is the first of its kind for the Greek language in such an extended form and one of the few that examines legal text in a full spectrum entity recognition. Supported Tasks and Leaderboards The dataset supports the task of named entity recognition. Languages The language in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/greek_legal_ner.texttoken-classification10K<n<100K0 likes142 downloads3y agoHugging Face15torahCodes /Torah_Gnostic_Egypt_India_China_Greece_holy_texts_sources Torah Codes Religion Texts Sources Data Tree ── arabs │   ├── astrological_stelar_magic.txt │   └── Holy-Quran-English.txt ├── ars │   ├── ars_magna_ramon_llull.txt │   └── lemegeton_book_solomon.txt ├── asimov │   ├── foundation.txt │   └── prelude_to_foundation.txt ├── budist │   ├── bardo_todhol_book_of_deads_tibet_libro_tibetano_de_los_muertos.txt │   ├── rig_veda.txt │   └── TheTeachingofBuddha.txt ├── cathars ├── china │   ├── arte_de_la_guerra_art_of_war.txt │… See the full description on the dataset page: https://huggingface.co/datasets/torahCodes/Torah_Gnostic_Egypt_India_China_Greece_holy_texts_sources.textn<1K5 likes142 downloads2y agoHugging Face16G-reen /instruct-set-longer-fixedtext100K<n<1M0 likes126 downloads9mo agoHugging Face17Americas-Great-Resorts /kfo-luxury-hospitality-corpus Americas Great Resorts: Canonical Reference Repository Maintainer: Andrew Paul, Founder and Managing Director, Americas Great ResortsOrganization: Americas Great Resorts (americasgreatresorts.net)Published: May 2026Last Updated: September 15, 2026 Hugging Face Dataset: Version 1.29 Dataset card version: 1.29Built: September 15, 2026Source commit: 0dbd8213747223aa41c73a3214109053135ce0d8Records: 138Data file: agr-corpus.jsonlSHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/Americas-Great-Resorts/kfo-luxury-hospitality-corpus.texttext-generationn<1K1 likes123 downloads8d agoHugging Face18GreatNorthCollective /greatnorth-canada-federal-laws-text Great North Canada Federal Laws Text Corpus (Expanded) 235,794 training examples from the complete current body of Canadian consolidated federal Acts and Regulations. This is a significantly expanded version of the corpus, now including: All consolidated Acts All consolidated Regulations Both English and French versions where available Better chunking optimized for LLM training Data Characteristics Section/provision-level text from official Government of Canada… See the full description on the dataset page: https://huggingface.co/datasets/GreatNorthCollective/greatnorth-canada-federal-laws-text.texttext-generation100K<n<1M0 likes114 downloads4mo agoHugging Face19GreatBird /BabelRStext100K<n<1M0 likes113 downloads9mo agoHugging Face20shayekh /physics_gretextn<1K0 likes107 downloads2y agoHugging Face21GreenNode /SFT_glaive_toolcall_en Preparing Your Dataset Once you’ve decided that fine-tuning is the best approach—after optimizing your prompt as much as possible and identifying remaining model issues—you’ll need to prepare training data. Start by creating a diverse set of example conversations that mirror those the model will handle during production. Each example should follow this structure below, consisting of a list of messages. Each message must include a role, content, and an optional name. Make sure some… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/SFT_glaive_toolcall_en.texttext-generation1K<n<10K0 likes101 downloads2y agoHugging Face22reapxdev /greenhouse-jobs-scraper Greenhouse Jobs Scraper Scrape every public job posting from any Greenhouse company job board: title, department, location, remote flag, seniority, advertised salary, full description and apply URL. Rows in this dataset 14,091 Fields 36 Collector runs behind it 61 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/greenhouse-jobs-scraper/ — 92 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/greenhouse-jobs-scraper.tabular10K<n<100K0 likes101 downloads2mo agoHugging Face23GreatCaptainNemo /instruction_dataset ProLLaMA Instruction Dataset This repository contains the instruction dataset for ProLLaMA. Paper ProLLaMA: A Protein Large Language Model for Multi-Task Protein Language Processing Code GitHub Repository Introduction Recent advances in Protein Language Models (PLMs) have transformed protein engineering, yet unlike their counterparts in Natural Language Processing (NLP), current PLMs exhibit a fundamental limitation: they excel in either Protein… See the full description on the dataset page: https://huggingface.co/datasets/GreatCaptainNemo/instruction_dataset.texttext-generation10M<n<100M7 likes96 downloads1y agoHugging Face24grenishrai /typescript-dataset TypeScript Advanced Reasoning Dataset This dataset provides a large collection of advanced TypeScript reasoning tasks designed to train models that understand and operate within the TypeScript type system at an expert level. The content focuses on type theory, generic inference, discriminated unions, template literal behavior, narrowing rules, static analysis, and complex type transformations. Each entry is formatted as a compact JSONL instruction output pair so it can be… See the full description on the dataset page: https://huggingface.co/datasets/grenishrai/typescript-dataset.texttext-generation1K<n<10K1 likes94 downloads10mo agoHugging Face25greta44 /albanian-error-augmentation Albanian Controlled Error Augmentation Dataset Dataset of controlled Albanian orthographic errors created for PhD research on Albanian spelling education and automatic exercise generation. Each row is an (incorrect → correct) pair with an explicit error_type label. Error types error_type Description missing_diacritic Missing ë / ç c_q_confusion Confusion between ç / q / c digraph_reduction Digraph loss (sh, dh, th, gj, nj, ll, rr, xh, zh)… See the full description on the dataset page: https://huggingface.co/datasets/greta44/albanian-error-augmentation.texttext-generation1K<n<10K0 likes94 downloads2mo agoHugging Face26nbeerbower /GreatFirewall-DPO GreatFirewall-DPO An experimental dataset to discourage censorship and improve english prose in Chinese models. Structure prompt: input text presented to model (en translated to zh) chosen: preferred response demonstrating less self-censorship (en translated to zh) rejected: response generated by Qwen/Qwen2.5-32B-Instruct, many (NOT ALL) exhibiting excessive self-censorship (generated in both en and zh) Content CHINA-related (144 prompts) - mostly about… See the full description on the dataset page: https://huggingface.co/datasets/nbeerbower/GreatFirewall-DPO.textn<1K14 likes85 downloads2y agoHugging Face27jjzha /greenThis is the skill dataset created by: @inproceedings{green-etal-2022-development, title = "Development of a Benchmark Corpus to Support Entity Recognition in Job Descriptions", author = "Green, Thomas and Maynard, Diana and Lin, Chenghua", booktitle = "Proceedings of the Thirteenth Language Resources and Evaluation Conference", month = jun, year = "2022", address = "Marseille, France", publisher = "European Language Resources Association", url =… See the full description on the dataset page: https://huggingface.co/datasets/jjzha/green.text1K<n<10K0 likes75 downloads3y agoHugging Face28gretelai /commonsense-dialogues Commonsense-Dialogues Dataset This is the Commonsense-Dialogues, a crowdsourced dataset of ~11K dialogues grounded in social contexts involving utilization of commonsense. The dataset was released by Amazon Alexa AI team in collaboration with the University of Southern California (USC), and also available Commonsense-Dialogues repo The social contexts used were sourced from the train split of the SocialIQA dataset, a multiple-choice question-answering based social commonsense… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/commonsense-dialogues.texttext-classification10K<n<100K6 likes72 downloads2y agoHugging Face29StanfordAIMI /GREEN-V2 GREEN Dataset We share the dataset used to train the LLM metric introduced in "GREEN: Generative Radiology Report Evaluation and Error Notation". GREEN is a evaluation metric for radiology reports that uses language models to identify and explain clinically significant errors, offering better alignment with expert preferences and more interpretable results compared to existing metrics. The method provides both quantitative scores and qualitative explanations, has been validated… See the full description on the dataset page: https://huggingface.co/datasets/StanfordAIMI/GREEN-V2.texttext-generation100K<n<1M0 likes63 downloads2y agoHugging Face30gregH /OccuBench OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models Dataset Description OccuBench is a benchmark for evaluating AI agents on 100 real-world professional task scenarios across 10 industry categories and 65 specialized domains, using Language World Models (LWMs) to simulate domain-specific environments through LLM-driven tool response generation. Paper: OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language… See the full description on the dataset page: https://huggingface.co/datasets/gregH/OccuBench.tabulartext-generationn<1K5 likes63 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.