datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GUI-World
GUI-World: A Dataset for GUI-Orientated Multimodal Large Language Models
Dataset: GUI-World
Overview
GUI-World introduces a comprehensive benchmark for evaluating MLLMs in dynamic and complex GUI environments. It features extensive annotations covering six GUI scenarios and eight types of GUI-oriented questions. The dataset assesses state-of-the-art ImageLLMs and VideoLLMs, highlighting their limitations in handling dynamic and multi-step tasks. It provides… See the full description on the dataset page: https://huggingface.co/datasets/ONE-Lab/GUI-World.guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/epfl-llm/guidelines.long-gui-tasks-v1
Long GUI Tasks HF Release v1
This directory is a local release package prepared for publishing the long-horizon GUI task dataset to Hugging Face.
Contents
metadata/train.jsonl: training split in ms-swift style messages + images format
metadata/test.jsonl: test split in the same format
metadata/train_manifest.jsonl: training sample manifest with metadata and image_ids
metadata/test_manifest.jsonl: test sample manifest
indexes/image_index.jsonl: unique image registry with… See the full description on the dataset page: https://huggingface.co/datasets/yzxjb/long-gui-tasks-v1.schema_guided_dstc8The Schema-Guided Dialogue dataset (SGD) was developed for the Dialogue State Tracking task of the Eights Dialogue Systems Technology Challenge (dstc8).
The SGD dataset consists of over 18k annotated multi-domain, task-oriented conversations between a human and a virtual assistant.
These conversations involve interactions with services and APIs spanning 17 domains, ranging from banks and events to media, calendar, travel, and weather.
For most of these domains, the SGD dataset contains multiple different APIs, many of which have overlapping functionalities but different interfaces,
which reflects common real-world scenarios.guiowl-curated-corpus
GUI-Owl Curated Corpus
This dataset publishes the full curated mobile GUI-agent supervised fine-tuning corpus in a unified norm1000 mobile_use action format. Each row pairs a mobile UI screenshot with an instruction and a normalized target tool call for training GUI agents.
The published files are the curated parquet shards as produced by the source canonicalizers. No parquet shards are merged, re-sharded, or sampled during upload.
Sources
Source
Episodes… See the full description on the dataset page: https://huggingface.co/datasets/KMK040412/guiowl-curated-corpus.GUIMid
Breaking the Data Barrier – Building GUI Agents Through Task Generalization
🐙 GitHub | 📝 Paper | 🤗 Mid-training Data | 🤗 Post-Training Data
TODO List
Report and release the GUIMid with larger size and more domains (10th May expecetd)
1. Data Overview
AgentBoard is composed of 9 diverse tasks: 7 vision and language tasks and 4 lanuage only tasks.
The performances of different domains as mid-training data are as follows:
Domains… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/GUIMid.go-swe-bench-v0
go_swe_bench v0 — real Go bug fixes, verified by the Go toolchain
246 tasks from 79 real Go repositories. Each task is a bug-fix commit whose co-committed test is red on the
parent and green on the fix. No LLM anywhere in the build.
Mined on 2026-09-19 from the GuildLM Go mining pipeline by inverting the filter that had thrown the tests
away (the pipeline was built for SFT data; a benchmark needs the opposite). Every task was verified twice
with go test: green at the commit (≥ 1… See the full description on the dataset page: https://huggingface.co/datasets/guildlm/go-swe-bench-v0.EC-Guide
This repo is only used for dataset viewer. Please download from here.
Amazon KDDCup 2024 Team ZJU-AI4H’s Solution and Dataset (Track 2 Top 2; Track 5 Top 5)
The Amazon KDD Cup’24 competition presents a unique challenge by focusing on the application of LLMs in E-commerce across multiple tasks. Our solution for addressing Tracks 2 and 5 involves a comprehensive pipeline encompassing dataset construction, instruction tuning, post-training quantization, and inference… See the full description on the dataset page: https://huggingface.co/datasets/AiMijie/EC-Guide.whatsapp-practical-guides
WhatsApp 实用指南合集(200 篇)
本仓库提供 200 篇中文实用教程,覆盖 WhatsApp、WhatsApp Web、WhatsApp Business 与 WhatsApp Business Platform 的常见操作与场景问题。
内容面向个人用户、远程办公者、小团队与客服人员,帮助读者更安全、更高效地完成登录、备份、协作与日常沟通。
每篇通常包含:背景说明、前置条件、逐步操作与完成标准、正误对照、可复制模板、边界情况、练习清单与 FAQ。
可按目录自学,也可作为团队内部培训材料使用。
官方源站
本合集涉及的产品入口,一律以以下官方源站为准:
👉 WhatsApp 网页版 —— 官方源站 🌐 https://web.whatsapp.com
👉 WhatsApp 帮助中心 —— 官方源站 🌐 https://faq.whatsapp.com
👉 WhatsApp Business —— 官方源站 🌐 https://www.whatsapp.com/business
👉 WhatsApp… See the full description on the dataset page: https://huggingface.co/datasets/fafafatt/whatsapp-practical-guides.EC-Guide
Amazon KDDCup 2024 Team ZJU-AI4H’s Solution and Dataset (Track 2 Top 2; Track 5 Top 5)
The Amazon KDD Cup’24 competition presents a unique challenge by focusing on the application of LLMs in E-commerce across multiple tasks. Our solution for addressing Tracks 2 and 5 involves a comprehensive pipeline encompassing dataset construction, instruction tuning, post-training quantization, and inference optimization. The core of our strategy is EC-Guide specifically tailored for E-commerce… See the full description on the dataset page: https://huggingface.co/datasets/AI4H/EC-Guide.Myanmar-Tuberculosis-Guidelines-Instructions
Myanmar Tuberculosis Guidelines Instructions
A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages.
Authors: Min Si Thu, Khin Myat Noe
Abstract
Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.japan-math-philosophy-prompts
Japan Math Philosophy Prompts
Microdataset autoral com problemas que combinam matemática e reflexão
filosófica. Há 24 registros: oito instâncias editoriais, cada uma localizada em
pt-BR, en e ja e mantida integralmente no split train.
Todo o conteúdo foi gerado por modelo e permanece sem revisão humana. As
respostas matemáticas funcionam como gabaritos curtos; os critérios filosóficos
indicam qualidades esperadas de uma justificativa, não uma opinião obrigatória.… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/japan-math-philosophy-prompts.mutopia_guitar_dataset
Mutopia Guitar Dataset
Dataset Summary
Mutopia guitar dataset consists of the soloist guitar pieces of the Mutopia Project. I encoded the MIDI files into text tokens using the excellent implementation of Dr. Tristan Beheren of the paper: MMM: Exploring Conditional Multi-Track Music Generation with the Transformer.
The dataset mainly contains guitar music from western classical composers, such as Sor, Aguado, Carcassi, and Giuliani.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/juancopi81/mutopia_guitar_dataset.task879_schema_guided_dstc8_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task879_schema_guided_dstc8_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task879_schema_guided_dstc8_classification.BazzBasic_AI_Guide
Dataset Card — BazzBasic AI Guide
The BazzBasic AI Guide is not a JSON dataset.
The BazzBasic AI Guide is a powerful guide "produced and approved" by Claude.ai, ChatGPT and Mistral Le Chat, specifically designed to be interpreted by modern AI.
Usage
Download latest BazzBasic-AI-guide from Files and versions
Upload it to your AI's prompt or project file and it will instantly become a BazzBasic expert.
Guide Summary
This guide contains the official… See the full description on the dataset page: https://huggingface.co/datasets/EkBass/BazzBasic_AI_Guide.gui_actor_webdataset
GUI-Actor WebDataset
A WebDataset format version of the GUI-Actor dataset for training vision-language models on GUI interaction tasks.
Usage
import webdataset as wds
# Load the dataset
dataset = wds.WebDataset("path/to/shards-*.tar")
dataset = dataset.decode("pilrgb").to_tuple("jpg", "json")
for image, metadata in dataset:
# Process image and metadata
pass
Citation
Please cite the original GUI-Actor paper if you use this dataset in your research.
guitar_tabDataset of music tablature, in alphaTex (https://alphatab.net/docs/alphatex)
format, converted from Guitar Pro files (gp3, gp4, gp5, which are downloaded
from https://rutracker.org/forum/viewtopic.php?t=2888130anthropic-Awareness-interview
anthropic-Awareness-interview
This dataset contains full transcripts of user research interviews where an AI assistant (Claude) interviews people about how they use AI in their work and how they feel about that collaboration.[web:1] Each example includes a long meta-cognitive system prompt plus a complete back-and-forth conversation.
Dataset overview
Domain: Human–AI interaction in professional and day-to-day work.
Format: Multi-turn chat logs with explicit roles.
Scale:… See the full description on the dataset page: https://huggingface.co/datasets/Guilherme34/anthropic-Awareness-interview.guidelinesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/epfl-llm/guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/guidelines.guia-de-colaboracao-do-ecossistema
Guia de Colaboração do Ecossistema de Inteligência Artificial — AI Brasil
Documento fundador da colaboração no ecossistema ai.eco.br · somos.aibrasil.ai · sou.aibrasil.ai
Este repositório publica o conteúdo do Guia de Colaboração da comunidade AI Brasil: a doutrina, a jornada de participação, o vocabulário de papéis, as regras do jogo, a camada prática da plataforma e o plano editorial da revista impressa de 48 páginas. É o material-base para quem quer entender como se colabora… See the full description on the dataset page: https://huggingface.co/datasets/Julianokimura/guia-de-colaboracao-do-ecossistema.guidelinesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/epfl-llm/guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-b1/guidelines.iceland-tech-christian-ethics-prompts
Fictional Icelandic Landscapes, Technology and Christian Ethics Prompts
This microdataset contains 24 original discussion prompts arranged as 12
parallel pt-BR/English pairs. Each explicitly fictional scenario combines a
landscape motif inspired by Iceland, a technology-governance dilemma, and
concepts that may be explored through Christian ethics. The records do not
describe real Icelandic institutions, policies, communities, or practices, and
they do not claim that Christians… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/iceland-tech-christian-ethics-prompts.HapticCapFull codes can be found in : https://github.com/LeMei/HapticCap
📌 HapticCap: A Multimodal Dataset and Task for Understanding User Experience of Vibration Haptic Signals
Arxiv: https://arxiv.org/pdf/2507.13318? (Findings of EMNLP2025)
📖 Introduction
HapticCap is a multimodal dataset and benchmark task designed for understanding user experience of vibration-based haptic signals.It provides a new resource for research at the intersection of haptics, text, and multimodal… See the full description on the dataset page: https://huggingface.co/datasets/GuiminHu/HapticCap.gui_grounding_dataset-1k
Supported Tasks
Natural Language → GUI Action Grounding
Convert user instructions into JSON action objects.
Instruction Following
Models learn to interpret varying natural language formulations (e.g., “press submit” vs “click the submit button”).
Multi-step UI Automation
Some samples involve sequences of actions (e.g., open site → type → press Enter → screenshot).
Languages
English (en)
Generated with simple variations (synonyms, phrasings).
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ArunKr/gui_grounding_dataset-1k.guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K… See the full description on the dataset page: https://huggingface.co/datasets/georgeqiao12138/guidelines.epfl-llm_guidelines_axolotl-completionepfl-llm/guidelines converted to work with axolotl completion or pretraining.
schema_guided_dstc8
Dataset Card for The Schema-Guided Dialogue Dataset
Dataset Summary
This dataset is a clone and adaptation of the original dataset google-research-datasets/schema_guided_dstc8 such that it is compatible with Dataset@4.0.0.
The Schema-Guided Dialogue dataset (SGD) was developed for the Dialogue State Tracking task of the Eights Dialogue Systems Technology Challenge (dstc8).
The SGD dataset consists of over 18k annotated multi-domain, task-oriented conversations between a… See the full description on the dataset page: https://huggingface.co/datasets/Deliangus/schema_guided_dstc8.br-sovereign-llm-corpus
BR Sovereign LLM Corpus
Status
This public repository is an audited corpus protocol and initial validation
snapshot. Public release 0.1.1 contains only a small, explicitly identified
corpus sample for pipeline verification. It is not a target-scale pretraining
corpus and does not support model-quality claims.
The scientific corpus remains under construction. Aggregate counts, source
shares, and token counts are not reported until a content-addressed snapshot… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/br-sovereign-llm-corpus.Sora-Ecommerce-Guide
Sora Ecommerce Guide Dataset
This dataset contains comprehensive documentation, user guides, admin operating procedures, and system flow architectures for the Sora Ecommerce platform, structured in flat instruction/input/output format matching standard fine-tuning benchmarks.
Splits
train: 9 samples
test: 2 samples
Features
instruction: System/task instruction context.
input: The prompt, question, or user query.
output: Complete step-by-step… See the full description on the dataset page: https://huggingface.co/datasets/HeinKoZin/Sora-Ecommerce-Guide.guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/WassimLab/guidelines.
