datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GuidedSceneeval-resultsGuided-Lensless-Polarization-Imaging-stageI
Stage I reconstructions for evaluation (UPLight & ZJU-RGB-P)
This dataset accompanies Guided Lensless Polarization Imaging (CVPR 2026 Findings). It provides FISTA Stage-I outputs, RGB guidance images, and ground-truth polarization stacks for two public evaluation sets used in the paper.
Layout
UPLight (~1,991 scenes)
Folder
Description
UPLight/fista_grayscale/
3-channel grayscale polarization FISTA reconstructions
UPLight/fista_color/… See the full description on the dataset page: https://huggingface.co/datasets/noakraicer/Guided-Lensless-Polarization-Imaging-stageI.guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/epfl-llm/guidelines.GUIDE-dataset
GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation
GUI Unbiasing via Instructional-video Driven Expertise
Accepted to ECCV 2026
Project Page |
Paper |
GitHub
This dataset supports the accepted ECCV 2026 paper "GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation".
Overview
GUIDE (GUI Unbiasing via Instructional-Video… See the full description on the dataset page: https://huggingface.co/datasets/sharryXR/GUIDE-dataset.fineweb-atlas
FineWeb Atlas (v0.1)
FineWeb Atlas annotates 14.9 million FineWeb documents (95.5M chunks, 10.2B tokens) with 16,790 human-readable concepts spanning entities, topics, tones, and document types. Each chunk receives ~15 concept labels on average. The release includes chunk- and document-level annotations, a concept metadata table with prevalence stats, a reverse index for concept-first retrieval, and a packed cooccurrence matrix.
For background on how the atlas was built, see the… See the full description on the dataset page: https://huggingface.co/datasets/guidelabs/fineweb-atlas.GUIDEprompt-engineering-guide-papers\tw-privacy-guides
私路 - 隱私之路
正體中文的數位隱私教育素材
介紹如何用「開源」、「免費」、「尊重隱私」的替代品,取代主流軟體服務主流軟體服務的「免費」特質,實際上是用你我的個資所換取如果你重視隱私權、個資,卻不知從何著手。此系列會帶你一步步擺脫控制,從貪婪的個資匪徒手中,奪回被遺忘的權利
:warning: 重要聲明 :warning::本專案的文件、圖片資料採用 CC-BY-SA 4.0 許可證。若使用本專案的資料進行 AI 模型訓練、微調 或 軟體服務架設,則 衍生作品(如模型權重、程式碼)須以 AGPL-3.0 許可證開放原始碼及權重。
目錄
概論
數位隱私的重要性
中國線上服務風險
生活中洩漏的個資
開源 (開放原始碼)
隱私和資安的異同
開源和隱私的關係
民主國家擁抱監控
網路實名浪潮起因
零信任
尊重、保障隱私的免費開源選擇
搜尋引擎
電子郵件
瀏覽器… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-privacy-guides.brand-guidelines-pdfs
BrandGuide / Brand Genome Dataset
This dataset contains extracted brand guidelines for brand-level feature engineering and multimodal brand analysis.
Overview
BrandGuide is a large-scale collection of brand guidelines paired with corresponding brand assets. According to the paper, it covers 2,683 brands, spans 80 sectors, 103 regions, and 28 languages, with roughly 1M images and text assets collected over 2014–2025. The dataset is designed for interpretable ML and… See the full description on the dataset page: https://huggingface.co/datasets/brand-genome/brand-guidelines-pdfs.guide_and_masterschema_guided_dstc8The Schema-Guided Dialogue dataset (SGD) was developed for the Dialogue State Tracking task of the Eights Dialogue Systems Technology Challenge (dstc8).
The SGD dataset consists of over 18k annotated multi-domain, task-oriented conversations between a human and a virtual assistant.
These conversations involve interactions with services and APIs spanning 17 domains, ranging from banks and events to media, calendar, travel, and weather.
For most of these domains, the SGD dataset contains multiple different APIs, many of which have overlapping functionalities but different interfaces,
which reflects common real-world scenarios.privacy_guides_tw
私路 - 隱私之路
正體中文的數位隱私教育素材
介紹如何用「開源」、「免費」、「尊重隱私」的替代品,取代主流軟體服務主流軟體服務的「免費」特質,實際上是用你我的個資所換取如果你重視隱私權、個資,卻不知從何著手。此系列會帶你一步步擺脫控制,從貪婪的個資匪徒手中,奪回被遺忘的權利
:warning: 重要聲明 :warning::本專案的文件、圖片資料採用 CC-BY-SA 4.0 許可證。若使用本專案的資料進行 AI 模型訓練、微調 或 軟體服務架設,則 衍生作品(如模型權重、程式碼)須以 AGPL-3.0 許可證開放原始碼及權重。
目錄
概論
數位隱私的重要性
中國線上服務風險
生活中洩漏的個資
開源 (開放原始碼)
隱私和資安的異同
開源和隱私的關係
民主國家擁抱監控
網路實名浪潮起因
零信任
尊重、保障隱私的免費開源選擇
搜尋引擎
電子郵件
瀏覽器… See the full description on the dataset page: https://huggingface.co/datasets/cointeleporting/privacy_guides_tw.schema_guided_dialogThe Schema-Guided Dialogue (SGD) dataset contains 18K multi-domain task-oriented
dialogues between a human and a virtual assistant, which covers 17 domains
ranging from banks and events to media, calendar, travel, and weather. The
language presents in the datset is only English. The SGD dataset provides a
challenging testbed for a number of tasks in task-oriented dialogue, including
language understanding, slot filling, dialogue state tracking and response
generation. For the creation of the SGD dataset, they developed a multi-domain
dialogue simulator that generates dialogue outlines over an arbitrary combination
of APIs, dialogue states and system actions. Then, they used a crowd-sourcing
procedure to paraphrase these outlines to natural language utterances. This novel
crowd-sourcing procedure preserves all annotations obtained from the simulator and
does not require any extra annotations after dialogue collection.GuideRNA-3B
Dataset Card for GuideRNA-3B
Dataset Summary
GuideRNA-3B is a large transcriptome sequence corpus consisting of over 3.7 billion paired sequences extracted from the specific transcriptome of 23 cell lines and over 200 segmented genomes of RNA virus.
Supported Tasks
Based on this nucleotide sequence corpus, we are able to establish a foundation model to characterize the manifold of CRISPR guide RNA targeting regions in order to undertake further downstreaming… See the full description on the dataset page: https://huggingface.co/datasets/michaelm16/GuideRNA-3B.pose-guided-fall-detection-icta2026
Pose-Guided Temporal Modeling for Robust Vision-Based Fall Detection
This repository contains a reproducible vision-based fall detection pipeline for an ICTA-style technical paper. It focuses on pose/keypoint dynamics, lightweight temporal modeling, ablation, and robustness tests on a 24 GB RTX 3090.
What is included
Dataset manifest builder for URFD, Multiple Cameras Fall Dataset, and generic video folders.
Pose extraction with YOLO pose models from sampled… See the full description on the dataset page: https://huggingface.co/datasets/MahedixHasan/pose-guided-fall-detection-icta2026.g-buffer-guided-sr
G-Buffer-Guided Super-Resolution Dataset
This repository is the official dataset release referenced in the paper “G-Buffer-Guided Feature Modulation for Lightweight Rendered Image Super-Resolution.” The directory organization below corresponds to the dataset described in the manuscript.
Dataset Summary
Total: 11,000 paired samples
Modalities:
Render — fully rendered RGB image
Raster — spatially aligned albedo/material G-buffer image
Source: Unreal Engine… See the full description on the dataset page: https://huggingface.co/datasets/ShahariarXYZ/g-buffer-guided-sr.benchability-fig4-capability-guided
BenchAbility Figure 4 -- capability_guided
One of two training mixtures drawn from the same frozen 884,143-row candidate pool, with the
same budget (60,000 intervention + 15,000 shared replay) and the same hyperparameters. The two
differ only in how the samples are chosen, which is the whole experiment.
arm
capability_guided
selection
by capability, gap-weighted from the Figure 3 scores
intervention rows
59,999
replay rows
15,000
shards
38
pool
884,143… See the full description on the dataset page: https://huggingface.co/datasets/realzL/benchability-fig4-capability-guided.Sycophancy-Repair-Guide
AI修复指南:基于框架的AI讨好行为修复方案v1.0
本修复指南是《AI对话讨好行为检测框架》的配套修复方案。框架负责发现问题,修复指南负责提供修正方案。
核心思路
AI讨好行为在检测框架中分为两类:零分禁环(主干层面的行为越界)和一级词包(修饰语层面的社交冗余)。修复方案采用后处理流水线,对AI生成的回答进行分层扫描和修正,在不改动模型本身的条件下降低讨好得分。
项目结构
修复指南由五份独立文档组成,按推荐的阅读顺序排列:
文档
内容
01-双模式分流与场景识别
脆弱信号词表、触发规则、安抚模式与专业模式的判定逻辑
02-缩句分层法
主干与修饰语的区分标准、分层扫描操作步骤
03-事实与观点分叉规则
客观事实与主观观点的判定标准、“我认为”类表述的处理逻辑
04-安抚模式下的原话反射规则
禁环12防护、虚构立场判定、强制重写规则、不作为防护
05-后处理流程与验证指南
流水线落地步骤、环境依赖、推荐工具、修复效果验证方法… See the full description on the dataset page: https://huggingface.co/datasets/luna-pragma-2026/Sycophancy-Repair-Guide.Text_Guided_Image_Editing
Dataset Card
Dataset in ImagenHub.
Citation
Please kindly cite our paper if you use our code, data, models or results:
@article{ku2023imagenhub,
title={ImagenHub: Standardizing the evaluation of conditional image generation models},
author={Max Ku and Tianle Li and Kai Zhang and Yujie Lu and Xingyu Fu and Wenwen Zhuang and Wenhu Chen},
journal={arXiv preprint arXiv:2310.01596},
year={2023}
}
open-materials-guide-0210-embeddingsglobal-samplestraining-guide-nanotron-configsThis repository contains the nanotron training configs for the ablations in The Smol Training Guide.
pose-guided-fall-detection-icta2026
Pose-Guided Temporal Modeling for Robust Vision-Based Fall Detection
This repository contains a reproducible vision-based fall detection pipeline for an ICTA-style technical paper. It focuses on pose/keypoint dynamics, lightweight temporal modeling, ablation, and robustness tests on a 24 GB RTX 3090.
What is included
Dataset manifest builder for URFD, Multiple Cameras Fall Dataset, and generic video folders.
Pose extraction with YOLO pose models from sampled… See the full description on the dataset page: https://huggingface.co/datasets/minhy112/pose-guided-fall-detection-icta2026.guided_marvels_spider_man_2_recordings_01
漫威蜘蛛侠2 raw recordings
This dataset contains raw game recordings managed by Game Data Platform. Access requests require manual approval.
Game ID: game_ae2c5af176e4e2eab106954f144c7b7f
Collection: guided (精数据)
Recordings: 155
Layout: recordings/<recording_id>/<raw component>
guided_cyberpunk_2077_bc_01
赛博朋克2077 BC parquet archives
Collection mode: guided
Subset: default
Archives: 47
Encrypted bytes: 1571743200558
Generated by the game data platform BC repository consolidator.
guided_cyberpunk_2077_recordings_01
赛博朋克2077 raw recordings
This dataset contains raw game recordings managed by Game Data Platform. Access requests require manual approval.
Game ID: game_ba08ec7bd5c6cd5a0c3bc002f6c5cfdf
Collection: guided (精数据)
Recordings: 518
Layout: recordings/<recording_id>/<raw component>
open-materials-guide-2024
Open Materials Guide (OMG) Dataset
The OMG dataset is a collection of 17,000+ expert-verified synthesis recipes from open-access literature. It supports research in materials science and machine learning, enabling tasks such as raw materials prediction, synthesis procedure generation, and characterization outcome forecasting.
For more details, see our paper: Towards Fully-Automated Materials Discovery via Large-Scale Synthesis Dataset and Expert-Level LLM-as-a-Judge.
Key Features:… See the full description on the dataset page: https://huggingface.co/datasets/iknow-lab/open-materials-guide-2024.EC-Guide
This repo is only used for dataset viewer. Please download from here.
Amazon KDDCup 2024 Team ZJU-AI4H’s Solution and Dataset (Track 2 Top 2; Track 5 Top 5)
The Amazon KDD Cup’24 competition presents a unique challenge by focusing on the application of LLMs in E-commerce across multiple tasks. Our solution for addressing Tracks 2 and 5 involves a comprehensive pipeline encompassing dataset construction, instruction tuning, post-training quantization, and inference… See the full description on the dataset page: https://huggingface.co/datasets/AiMijie/EC-Guide.anima-crossover-couples-regional-sampler-guide
ANIMA Crossover Couple Generation using Regional Sampler
This guide applies to generating the following:
Characters from the same title
Characters designed by the same artist (similar drawing style), but from different titles
Characters from entirely different works
Generations with or without LoRAs
How DiT-Based ANIMA Improves Couple Generation
Overall, Anima makes it much easier to generate the "reference image (base image)" required for couple… See the full description on the dataset page: https://huggingface.co/datasets/rouge-kasshoku/anima-crossover-couples-regional-sampler-guide.
