datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PADBench
Vincent-HKUSTGZ/PADBench
This repository contains the complete dataset from the finish_transfer_rank folder, including Llama2-7B models fine-tuned with PEFTGuard for different datasets.
Content Structure
This dataset includes the following subdirectories:
.cache/: Contains model files for .cache dataset
chatglm6b_toxic_backdoors_hard_rank256_qv/: Contains model files for chatglm6b_toxic_backdoors_hard_rank256_qv dataset
flan_t5_xl_toxic_backdoors_hard_rank256_qv/:… See the full description on the dataset page: https://huggingface.co/datasets/Vincent-HKUSTGZ/PADBench.PaddleOCR-VL_demoReal5-OmniDocBench
Real5-OmniDocBench
A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild
Leaderboard | Overview | Dataset | Evaluation | Submit Results | Citation
Real5-OmniDocBench measures the robustness of document parsing systems under five physical acquisition conditions: Scanning, Warping, Screen-Photography, Illumination, and Skew. It reconstructs the same 1,355 pages from OmniDocBench v1.5 in every condition, producing 6,775 images in total. The one-to-one… See the full description on the dataset page: https://huggingface.co/datasets/PaddlePaddle/Real5-OmniDocBench.webfaq-retrievalWebFAQ Retrieval Dataset
Overview |
Details |
Structure |
Examples |
Considerations |
License |
Citation |
Contact |
Acknowledgement
Overview
The WebFAQ Retrieval Dataset is a carefully filtered and curated subset of the broader WebFAQ Q&A Dataset.It is purpose-built for Information Retrieval (IR) tasks, such as training and evaluating dense or sparse retrieval models in multiple languages.
Each of the… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq-retrieval.berita-padupad-ufes-20
PAD-UFES-20
This dataset repository mirrors PAD-UFES-20 for reproducible teledermatology
experiments in mlops-teledermatology.
PAD-UFES-20 contains smartphone clinical images of skin lesions plus tabular
metadata. The dataset includes 2,298 images from 1,373 patients and 1,641 skin
lesions. The labels used by this project are:
ACK: Actinic keratosis
BCC: Basal cell carcinoma
MEL: Melanoma
NEV: Nevus
SCC: Squamous cell carcinoma, including Bowen's disease/SCC in situ
SEK: Seborrheic… See the full description on the dataset page: https://huggingface.co/datasets/SalmaneExploring/pad-ufes-20.CoRECoRE: Controlled Retrieval Evaluation Dataset
Motivation |
Dataset Overview |
Dataset Construction |
Dataset Structure |
Qrels Format |
Evaluation |
Citation |
Links |
Contact
CoRE (Controlled Retrieval Evaluation) is a benchmark dataset designed for the rigorous evaluation of embedding compression techniques in information retrieval.
🔍 Motivation
Embedding compression is essential for scaling… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/CoRE.webfaqWebFAQ Q&A Dataset
Overview |
Details |
Structure |
Examples |
Considerations |
License |
Citation |
Contact |
Acknowledgement
Overview
The WebFAQ Q&A Dataset is a broad-coverage corpus of 96 million natural question-answer (QA) pairs in 75 languages, gathered from FAQ pages on the web. It leverages structured schema.org FAQPage annotations, making it a unique resource for large-scale Question Answering… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq.PadChest-GR-Xraywebfaq-v2WebFAQ 2.0: Multilingual FAQ Q&A Dataset with Hard Negatives
Overview |
What's New in v2.0 |
Dataset Statistics |
Structure |
Bitext Alignments |
Hard Negatives |
Training Strategies |
Examples |
Considerations |
License |
Citation |
Contact
Overview
WebFAQ 2.0 is a large-scale multilingual dataset of 198 million natural question–answer pairsacross 108 languages, mined from structured… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq-v2.RoboTwin_move_stapler_pad_randomizedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 500,
"total_frames": 73633,
"total_tasks": 500,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/suz22/RoboTwin_move_stapler_pad_randomized.RoboTwin_move_pillbottle_pad_randomizedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 500,
"total_frames": 70667,
"total_tasks": 474,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/suz22/RoboTwin_move_pillbottle_pad_randomized.RoboTwin_place_mouse_pad_randomizedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 500,
"total_frames": 73008,
"total_tasks": 498,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/suz22/RoboTwin_place_mouse_pad_randomized.webfaq-bitextsWebFAQ Bilingual Datasets (Bitexts)
Overview |
Details |
Structure |
Examples |
Considerations |
License |
Citation |
Contact |
Acknowledgement
Overview
The WebFAQ Bilingual Datasets (a.k.a. Bitexts) are derived from the WebFAQ Q&A Dataset, but instead of monolingual question-answer (QA) pairs, each entry here contains aligned QA pairs in two different languages. These alignments are created via… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq-bitexts.padma-meghna-riverbank-erosionlm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private
Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v1.0
Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v1.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private.pad-us-3
United States Protected Areas Database
U.S. Geological Survey (USGS) Gap Analysis Project (GAP), 2022, Protected Areas Database of the United States (PAD-US) 3.0: U.S. Geological Survey data release, https://doi.org/10.5066/P9Q9LQ4B.
The PAD database classifications are complex and include overlapping and apparently duplicate polygons. For instance, national park boundaries can be found are listed in Proclamation feature class and also in Fee class (showing internal holes within… See the full description on the dataset page: https://huggingface.co/datasets/boettiger-lab/pad-us-3.pi0_conversion_no_padThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "panda",
"total_episodes": 1417,
"total_frames": 166855,
"total_tasks": 33,
"total_videos": 0,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:1417"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mlfu7/pi0_conversion_no_pad.moltbook-corpus
Moltbook Corpus: Agent Social Behavior Dataset
This dataset provides for the research paper:
"FORM WITHOUT FUNCTION: AGENT SOCIAL BEHAVIOR IN THE MOLTBOOK NETWORK"📄 https://arxiv.org/abs/2604.13052
📊 Dataset Statistics
Category
Total Count
Collection Period
Jan 27, 2026 – Mar 8, 2026
Total Posts
1,312,238
Total Comments
6,691,460
Total Profiles
120,811
Total Submolts Metadata
108,490
Annotated Posts
394,221
Annotated Comments
2,100,589… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/moltbook-corpus.pad3-image
PAD3: Multi-Domain Image Classification Dataset for Kids' Safety
Dataset Summary
The PAD3 (Protected Access Defense - Domain Detection) dataset is a curated collection of over 50,000 images designed specifically to train and evaluate computer vision models for child-safe content moderation. The dataset provides a robust framework for binary classification (Safe vs. Unsafe) and granular violation detection across multiple sensitive domains.
Focusing on high-risk categories… See the full description on the dataset page: https://huggingface.co/datasets/arkananta27/pad3-image.colab-paddle-env-for-ocrtts_dataset_1.05M_paddedfineweb-edu-dedup-train-5B-by-Llama-3.2-3B-tokenizer-2048-pack-padRefCOCOPatch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
[🔗 Released Code]
[🤗 Datasets] [🤗 Checkpoints]
[📄 Tech Report] [🤗 Paper]
Figure A. PaDT pipeline.
🌟 Introduction
We are pleased to introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables multimodal large language models (MLLMs) to directly generate both textual and visual outputs.At the core of PaDT are Visual Reference Tokens (VRTs). Unlike conventional MLLMs that represent… See the full description on the dataset page: https://huggingface.co/datasets/PaDT-MLLM/RefCOCO.lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private
Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v3.0
Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v3.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private.put_bottle_on_pad_25_07_23_parquettts_dataset_1.05M_padded_text_labels_oncubemaps_padding_16px_captioned_40kasd-7-tokenized-paddedasd-4-tokenized-padded
