datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/epfl-llm/guidelines.go-swe-bench-v0
go_swe_bench v0 — real Go bug fixes, verified by the Go toolchain
246 tasks from 79 real Go repositories. Each task is a bug-fix commit whose co-committed test is red on the
parent and green on the fix. No LLM anywhere in the build.
Mined on 2026-09-19 from the GuildLM Go mining pipeline by inverting the filter that had thrown the tests
away (the pipeline was built for SFT data; a benchmark needs the opposite). Every task was verified twice
with go test: green at the commit (≥ 1… See the full description on the dataset page: https://huggingface.co/datasets/guildlm/go-swe-bench-v0.guiowl-curated-corpus
GUI-Owl Curated Corpus
This dataset publishes the full curated mobile GUI-agent supervised fine-tuning corpus in a unified norm1000 mobile_use action format. Each row pairs a mobile UI screenshot with an instruction and a normalized target tool call for training GUI agents.
The published files are the curated parquet shards as produced by the source canonicalizers. No parquet shards are merged, re-sharded, or sampled during upload.
Sources
Source
Episodes… See the full description on the dataset page: https://huggingface.co/datasets/KMK040412/guiowl-curated-corpus.GUIMid
Breaking the Data Barrier – Building GUI Agents Through Task Generalization
🐙 GitHub | 📝 Paper | 🤗 Mid-training Data | 🤗 Post-Training Data
TODO List
Report and release the GUIMid with larger size and more domains (10th May expecetd)
1. Data Overview
AgentBoard is composed of 9 diverse tasks: 7 vision and language tasks and 4 lanuage only tasks.
The performances of different domains as mid-training data are as follows:
Domains… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/GUIMid.EC-Guide
This repo is only used for dataset viewer. Please download from here.
Amazon KDDCup 2024 Team ZJU-AI4H’s Solution and Dataset (Track 2 Top 2; Track 5 Top 5)
The Amazon KDD Cup’24 competition presents a unique challenge by focusing on the application of LLMs in E-commerce across multiple tasks. Our solution for addressing Tracks 2 and 5 involves a comprehensive pipeline encompassing dataset construction, instruction tuning, post-training quantization, and inference… See the full description on the dataset page: https://huggingface.co/datasets/AiMijie/EC-Guide.guidelinesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/epfl-llm/guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/guidelines.guidelinesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/epfl-llm/guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-b1/guidelines.Myanmar-Tuberculosis-Guidelines-Instructions
Myanmar Tuberculosis Guidelines Instructions
A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages.
Authors: Min Si Thu, Khin Myat Noe
Abstract
Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.japan-math-philosophy-prompts
Japan Math Philosophy Prompts
Microdataset autoral com problemas que combinam matemática e reflexão
filosófica. Há 24 registros: oito instâncias editoriais, cada uma localizada em
pt-BR, en e ja e mantida integralmente no split train.
Todo o conteúdo foi gerado por modelo e permanece sem revisão humana. As
respostas matemáticas funcionam como gabaritos curtos; os critérios filosóficos
indicam qualidades esperadas de uma justificativa, não uma opinião obrigatória.… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/japan-math-philosophy-prompts.mutopia_guitar_dataset
Mutopia Guitar Dataset
Dataset Summary
Mutopia guitar dataset consists of the soloist guitar pieces of the Mutopia Project. I encoded the MIDI files into text tokens using the excellent implementation of Dr. Tristan Beheren of the paper: MMM: Exploring Conditional Multi-Track Music Generation with the Transformer.
The dataset mainly contains guitar music from western classical composers, such as Sor, Aguado, Carcassi, and Giuliani.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/juancopi81/mutopia_guitar_dataset.task879_schema_guided_dstc8_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task879_schema_guided_dstc8_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task879_schema_guided_dstc8_classification.guitar_tabDataset of music tablature, in alphaTex (https://alphatab.net/docs/alphatex)
format, converted from Guitar Pro files (gp3, gp4, gp5, which are downloaded
from https://rutracker.org/forum/viewtopic.php?t=2888130gui_actor_webdataset
GUI-Actor WebDataset
A WebDataset format version of the GUI-Actor dataset for training vision-language models on GUI interaction tasks.
Usage
import webdataset as wds
# Load the dataset
dataset = wds.WebDataset("path/to/shards-*.tar")
dataset = dataset.decode("pilrgb").to_tuple("jpg", "json")
for image, metadata in dataset:
# Process image and metadata
pass
Citation
Please cite the original GUI-Actor paper if you use this dataset in your research.
anthropic-Awareness-interview
anthropic-Awareness-interview
This dataset contains full transcripts of user research interviews where an AI assistant (Claude) interviews people about how they use AI in their work and how they feel about that collaboration.[web:1] Each example includes a long meta-cognitive system prompt plus a complete back-and-forth conversation.
Dataset overview
Domain: Human–AI interaction in professional and day-to-day work.
Format: Multi-turn chat logs with explicit roles.
Scale:… See the full description on the dataset page: https://huggingface.co/datasets/Guilherme34/anthropic-Awareness-interview.Sora-Ecommerce-Guide
Sora Ecommerce Guide Dataset
This dataset contains comprehensive documentation, user guides, admin operating procedures, and system flow architectures for the Sora Ecommerce platform, structured in flat instruction/input/output format matching standard fine-tuning benchmarks.
Splits
train: 9 samples
test: 2 samples
Features
instruction: System/task instruction context.
input: The prompt, question, or user query.
output: Complete step-by-step… See the full description on the dataset page: https://huggingface.co/datasets/HeinKoZin/Sora-Ecommerce-Guide.guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K… See the full description on the dataset page: https://huggingface.co/datasets/georgeqiao12138/guidelines.iceland-tech-christian-ethics-prompts
Fictional Icelandic Landscapes, Technology and Christian Ethics Prompts
This microdataset contains 24 original discussion prompts arranged as 12
parallel pt-BR/English pairs. Each explicitly fictional scenario combines a
landscape motif inspired by Iceland, a technology-governance dilemma, and
concepts that may be explored through Christian ethics. The records do not
describe real Icelandic institutions, policies, communities, or practices, and
they do not claim that Christians… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/iceland-tech-christian-ethics-prompts.gui_grounding_dataset-1k
Supported Tasks
Natural Language → GUI Action Grounding
Convert user instructions into JSON action objects.
Instruction Following
Models learn to interpret varying natural language formulations (e.g., “press submit” vs “click the submit button”).
Multi-step UI Automation
Some samples involve sequences of actions (e.g., open site → type → press Enter → screenshot).
Languages
English (en)
Generated with simple variations (synonyms, phrasings).
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ArunKr/gui_grounding_dataset-1k.epfl-llm_guidelines_axolotl-completionepfl-llm/guidelines converted to work with axolotl completion or pretraining.
br-sovereign-llm-corpus
BR Sovereign LLM Corpus
Status
This public repository is an audited corpus protocol and initial validation
snapshot. Public release 0.1.1 contains only a small, explicitly identified
corpus sample for pipeline verification. It is not a target-scale pretraining
corpus and does not support model-quality claims.
The scientific corpus remains under construction. Aggregate counts, source
shares, and token counts are not reported until a content-addressed snapshot… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/br-sovereign-llm-corpus.gui_grounding_dataset-100
Supported Tasks
Natural Language → GUI Action Grounding
Convert user instructions into JSON action objects.
Instruction Following
Models learn to interpret varying natural language formulations (e.g., “press submit” vs “click the submit button”).
Multi-step UI Automation
Some samples involve sequences of actions (e.g., open site → type → press Enter → screenshot).
Languages
English (en)
Generated with simple variations (synonyms, phrasings).
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ArunKr/gui_grounding_dataset-100.guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/WassimLab/guidelines.paragen-security-sft-alpaca
paragen-security-sft-alpaca
Alpaca-format instruction-tuning data used to train the Vanilla security
baseline (and as the source for the tokenized multi-stream cache used to
train the Stream(Ours) security checkpoint) in the paragen_llm
security/prompt-injection-robustness experiments (Table 3: TensorTrust,
Gandalf, Purple, RuLES, StruQ, NESSiE, IFEval).
Format: JSONL, one object per line, fields instruction / input / output
(standard Alpaca schema).
Size: 48,538 examples.… See the full description on the dataset page: https://huggingface.co/datasets/guinansu/paragen-security-sft-alpaca.guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/Billy6310/guidelines.smolified-bengali-local-food-guide
🤏 smolified-bengali-local-food-guide
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-bengali-local-food-guide.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 638d3b25)
Records: 1050
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
prepware_study_guide-dataset
Prepware_Study_Guide Dataset
Generated by DocParserEngine.
Field
Value
Documents
1
Records
1
Schema
full
Usage
from datasets import load_dataset
ds = load_dataset("Remixonwin/prepware_study_guide-dataset")
guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/mmrech/guidelines.hulk_dataset_0.1This dataset is AFAIK (12 january 2024) the biggest ready to use open source dataset to finetune LLMs. It contains more than 3.8 million chat samples.
Its a collection of multiple different datasets. Some of them have been built using GPT4 or using scraped data. Here is the list:
gathnex/Gath_baize
teknium/openhermes
nomic-ai/gpt4all-j-prompt-generations
teknium/dataforge-economics
Anthropic/hh-rlhf: we kept only the selected prompts
teknium1_GPTeacher_codegen… See the full description on the dataset page: https://huggingface.co/datasets/guigux/hulk_dataset_0.1.epfl-llm_guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/minsu/epfl-llm_guidelines.gui-vs-cli
GUI-vs-CLI: A Unified Benchmark
This dataset contains task descriptions and verification specifications for 440 desktop software tasks from the GUI-vs-CLI benchmark.
The Hugging Face dataset is intended for browsing and lightweight programmatic access to task descriptions. Full runnable task assets, environment files, and execution code are maintained in the GitHub repository.
Files
data/tasks.jsonl: one task per line.
Fields
Each row contains:… See the full description on the dataset page: https://huggingface.co/datasets/rebeccazzzz/gui-vs-cli.
