datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
leetcodesymptom_to_diagnosis
Dataset Summary
This dataset contains natural language descriptions of symptoms labeled with 22 corresponding diagnoses. Gretel/symptom_to_diagnosis provides 1065 symptom descriptions in the English language labeled with 22 diagnoses, focusing on fine-grained single-domain diagnosis.
Data Fields
Each row contains the following fields:
input_text : A string field containing symptoms
output_text : A string field containing a diagnosis
Example:
{
"output_text": "drug… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/symptom_to_diagnosis.instruct-setbig_setars-magna-greatest-hits
Ars Magna Greatest Hits
The funniest and most apt anagrams of people, companies, products, titles, places and phrases, found by Ars Magna and kept by hand.
Every row is a real anagram: the words use exactly the input's letters, checked against a pinned revision of English OpenList (368bf0e4460461c985fca8bde49e4062d56c1516), and every word is in the tier the row names. Accented letters fold to their base letter, so Beyoncé has three e's. Nothing typed is ever replaced by… See the full description on the dataset page: https://huggingface.co/datasets/ryanjosephkamp/ars-magna-greatest-hits.instruct-set-longerproduct-taxonomy-bench
Dataset Summary
product-taxonomy-bench is an anonymised benchmark dataset for predicting Shopify Product Taxonomy categories from Shopify product tags.
This dataset does not include raw product titles, raw tags, or product URLs. Tags are anonymised as tagNNNNNN.
Start Here
Read this dataset card for the snapshot layout and field definitions.
Open the benchmark notebook at notebooks/product_taxonomy_bench.ipynb. It defaults to the fixed paper snapshot at revision… See the full description on the dataset page: https://huggingface.co/datasets/gregb/product-taxonomy-bench.Greetings-hi-for-train-Msh-v2Simple English greeting & everyday conversation dataset, built for training Msh.
Format
Each line is a JSON object:
{"user": "...", "ai": "..."}
Content
Basic greetings, small talk, farewells
Basic math, poems, short fairy tales
Basic code snippets (multiple languages)
Identity, ethics, general knowledge
ty
RoleBreak
RoleBreak
A benchmark for long-horizon role-playing robustness in spoken dialogue.
RoleBreak holds 310 roles, 6,688 human-verified user turns (21.6 per
conversation) and 11,743 fine-grained pass/fail criteria. Each conversation puts
a speech-to-speech model in character and then stresses it as context
accumulates — context-dependent probes and targeted interventions against role
consistency, interaction quality, safety, and affect.
This repository includes three things:
The… See the full description on the dataset page: https://huggingface.co/datasets/Greenbean/RoleBreak.medium_setgreenhouse-sensor-data
Pomona Greenhouse Sensor Data
This public research dataset contains greenhouse time-series files and Pomona
training-oriented JSONL derived from greenhouse sensor data. It supports
experiments in compact agricultural reasoners, digital twins, anomaly review,
and structured decision-support models.
Research data, not an operational control policy. Sensor records may be
incomplete, noisy, synthetic, transformed, or facility-specific. Do not use
dataset rows as direct actuator… See the full description on the dataset page: https://huggingface.co/datasets/Okyanus/greenhouse-sensor-data.amc12-full
AMC12 Dataset (Research-Oriented)
A structured dataset derived from the AMC 12 (American Mathematics Competitions), designed for LLM training, evaluation, and reinforcement learning (RL) on mathematical reasoning tasks.
This repository contains all AMC 12 problems from 2000–2025, making it one of the most complete AMC12 datasets available for research.
📘 Introduction
The AMC 12 is a 25-question, 75-minute multiple-choice examination aimed at high school… See the full description on the dataset page: https://huggingface.co/datasets/greenstainedglass/amc12-full.fable-5-coding-and-debugging-traces-synthetic-corrections
Model Synthetic Corrections
1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset.
Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.greek_legal_ner
Dataset Card for Greek Legal Named Entity Recognition
Dataset Summary
This dataset contains an annotated corpus for named entity recognition in Greek legislations. It is the first of its kind for the Greek language in such an extended form and one of the few that examines legal text in a full spectrum entity recognition.
Supported Tasks and Leaderboards
The dataset supports the task of named entity recognition.
Languages
The language in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/greek_legal_ner.Torah_Gnostic_Egypt_India_China_Greece_holy_texts_sources
Torah Codes Religion Texts Sources
Data Tree
── arabs
│ ├── astrological_stelar_magic.txt
│ └── Holy-Quran-English.txt
├── ars
│ ├── ars_magna_ramon_llull.txt
│ └── lemegeton_book_solomon.txt
├── asimov
│ ├── foundation.txt
│ └── prelude_to_foundation.txt
├── budist
│ ├── bardo_todhol_book_of_deads_tibet_libro_tibetano_de_los_muertos.txt
│ ├── rig_veda.txt
│ └── TheTeachingofBuddha.txt
├── cathars
├── china
│ ├── arte_de_la_guerra_art_of_war.txt
│… See the full description on the dataset page: https://huggingface.co/datasets/torahCodes/Torah_Gnostic_Egypt_India_China_Greece_holy_texts_sources.instruct-set-longer-fixedkfo-luxury-hospitality-corpus
Americas Great Resorts: Canonical Reference Repository
Maintainer: Andrew Paul, Founder and Managing Director, Americas Great ResortsOrganization: Americas Great Resorts (americasgreatresorts.net)Published: May 2026Last Updated: September 15, 2026
Hugging Face Dataset: Version 1.29
Dataset card version: 1.29Built: September 15, 2026Source commit: 0dbd8213747223aa41c73a3214109053135ce0d8Records: 138Data file: agr-corpus.jsonlSHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/Americas-Great-Resorts/kfo-luxury-hospitality-corpus.greatnorth-canada-federal-laws-text
Great North Canada Federal Laws Text Corpus (Expanded)
235,794 training examples from the complete current body of Canadian consolidated federal Acts and Regulations.
This is a significantly expanded version of the corpus, now including:
All consolidated Acts
All consolidated Regulations
Both English and French versions where available
Better chunking optimized for LLM training
Data Characteristics
Section/provision-level text from official Government of Canada… See the full description on the dataset page: https://huggingface.co/datasets/GreatNorthCollective/greatnorth-canada-federal-laws-text.BabelRSphysics_greSFT_glaive_toolcall_en
Preparing Your Dataset
Once you’ve decided that fine-tuning is the best approach—after optimizing your prompt as much as possible and identifying remaining model issues—you’ll need to prepare training data. Start by creating a diverse set of example conversations that mirror those the model will handle during production.
Each example should follow this structure below, consisting of a list of messages. Each message must include a role, content, and an optional name. Make sure some… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/SFT_glaive_toolcall_en.greenhouse-jobs-scraper
Greenhouse Jobs Scraper
Scrape every public job posting from any Greenhouse company job board: title, department, location, remote flag, seniority, advertised salary, full description and apply URL.
Rows in this dataset
14,091
Fields
36
Collector runs behind it
61
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/greenhouse-jobs-scraper/ — 92 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/greenhouse-jobs-scraper.instruction_dataset
ProLLaMA Instruction Dataset
This repository contains the instruction dataset for ProLLaMA.
Paper
ProLLaMA: A Protein Large Language Model for Multi-Task Protein Language Processing
Code
GitHub Repository
Introduction
Recent advances in Protein Language Models (PLMs) have transformed protein engineering, yet unlike their counterparts in Natural Language Processing (NLP), current PLMs exhibit a fundamental limitation: they excel in either Protein… See the full description on the dataset page: https://huggingface.co/datasets/GreatCaptainNemo/instruction_dataset.typescript-dataset
TypeScript Advanced Reasoning Dataset
This dataset provides a large collection of advanced TypeScript reasoning tasks designed to train models that understand and operate within the TypeScript type system at an expert level. The content focuses on type theory, generic inference, discriminated unions, template literal behavior, narrowing rules, static analysis, and complex type transformations.
Each entry is formatted as a compact JSONL instruction output pair so it can be… See the full description on the dataset page: https://huggingface.co/datasets/grenishrai/typescript-dataset.albanian-error-augmentation
Albanian Controlled Error Augmentation Dataset
Dataset of controlled Albanian orthographic errors created for PhD research on Albanian spelling education and automatic exercise generation.
Each row is an (incorrect → correct) pair with an explicit error_type label.
Error types
error_type
Description
missing_diacritic
Missing ë / ç
c_q_confusion
Confusion between ç / q / c
digraph_reduction
Digraph loss (sh, dh, th, gj, nj, ll, rr, xh, zh)… See the full description on the dataset page: https://huggingface.co/datasets/greta44/albanian-error-augmentation.GreatFirewall-DPO
GreatFirewall-DPO
An experimental dataset to discourage censorship and improve english prose in Chinese models.
Structure
prompt: input text presented to model (en translated to zh)
chosen: preferred response demonstrating less self-censorship (en translated to zh)
rejected: response generated by Qwen/Qwen2.5-32B-Instruct, many (NOT ALL) exhibiting excessive self-censorship (generated in both en and zh)
Content
CHINA-related (144 prompts) - mostly about… See the full description on the dataset page: https://huggingface.co/datasets/nbeerbower/GreatFirewall-DPO.greenThis is the skill dataset created by:
@inproceedings{green-etal-2022-development,
title = "Development of a Benchmark Corpus to Support Entity Recognition in Job Descriptions",
author = "Green, Thomas and
Maynard, Diana and
Lin, Chenghua",
booktitle = "Proceedings of the Thirteenth Language Resources and Evaluation Conference",
month = jun,
year = "2022",
address = "Marseille, France",
publisher = "European Language Resources Association",
url =… See the full description on the dataset page: https://huggingface.co/datasets/jjzha/green.commonsense-dialogues
Commonsense-Dialogues Dataset
This is the Commonsense-Dialogues, a crowdsourced dataset of ~11K dialogues grounded in social contexts involving utilization of commonsense. The dataset was released by Amazon Alexa AI team in collaboration with the University of Southern California (USC), and also available Commonsense-Dialogues repo
The social contexts used were sourced from the train split of the SocialIQA dataset, a multiple-choice question-answering based social commonsense… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/commonsense-dialogues.GREEN-V2
GREEN Dataset
We share the dataset used to train the LLM metric introduced in "GREEN: Generative Radiology Report Evaluation and Error Notation".
GREEN is a evaluation metric for radiology reports that uses language models to identify and explain clinically significant errors, offering better alignment with expert preferences and more interpretable results compared to existing metrics. The method provides both quantitative scores and qualitative explanations, has been validated… See the full description on the dataset page: https://huggingface.co/datasets/StanfordAIMI/GREEN-V2.OccuBench
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models
Dataset Description
OccuBench is a benchmark for evaluating AI agents on 100 real-world professional task scenarios across 10 industry categories and 65 specialized domains, using Language World Models (LWMs) to simulate domain-specific environments through LLM-driven tool response generation.
Paper: OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language… See the full description on the dataset page: https://huggingface.co/datasets/gregH/OccuBench.
