datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
webfaq-retrievalWebFAQ Retrieval Dataset
Overview |
Details |
Structure |
Examples |
Considerations |
License |
Citation |
Contact |
Acknowledgement
Overview
The WebFAQ Retrieval Dataset is a carefully filtered and curated subset of the broader WebFAQ Q&A Dataset.It is purpose-built for Information Retrieval (IR) tasks, such as training and evaluating dense or sparse retrieval models in multiple languages.
Each of the… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq-retrieval.CoRECoRE: Controlled Retrieval Evaluation Dataset
Motivation |
Dataset Overview |
Dataset Construction |
Dataset Structure |
Qrels Format |
Evaluation |
Citation |
Links |
Contact
CoRE (Controlled Retrieval Evaluation) is a benchmark dataset designed for the rigorous evaluation of embedding compression techniques in information retrieval.
🔍 Motivation
Embedding compression is essential for scaling… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/CoRE.webfaqWebFAQ Q&A Dataset
Overview |
Details |
Structure |
Examples |
Considerations |
License |
Citation |
Contact |
Acknowledgement
Overview
The WebFAQ Q&A Dataset is a broad-coverage corpus of 96 million natural question-answer (QA) pairs in 75 languages, gathered from FAQ pages on the web. It leverages structured schema.org FAQPage annotations, making it a unique resource for large-scale Question Answering… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq.webfaq-bitextsWebFAQ Bilingual Datasets (Bitexts)
Overview |
Details |
Structure |
Examples |
Considerations |
License |
Citation |
Contact |
Acknowledgement
Overview
The WebFAQ Bilingual Datasets (a.k.a. Bitexts) are derived from the WebFAQ Q&A Dataset, but instead of monolingual question-answer (QA) pairs, each entry here contains aligned QA pairs in two different languages. These alignments are created via… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq-bitexts.lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private
Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v1.0
Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v1.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private.RefCOCOPatch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
[🔗 Released Code]
[🤗 Datasets] [🤗 Checkpoints]
[📄 Tech Report] [🤗 Paper]
Figure A. PaDT pipeline.
🌟 Introduction
We are pleased to introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables multimodal large language models (MLLMs) to directly generate both textual and visual outputs.At the core of PaDT are Visual Reference Tokens (VRTs). Unlike conventional MLLMs that represent… See the full description on the dataset page: https://huggingface.co/datasets/PaDT-MLLM/RefCOCO.lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private
Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v3.0
Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v3.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private.pad-auto-solver-reviewed
PAD Reviewed Dataset
Canonical reviewed PAD board/orb artifacts for dw-indie/pad-auto-solver-reviewed. This repository
contains immutable reviewed package revisions and does not contain raw captures,
training runs, checkpoints, or model binaries.
Packages exported: 28
Active catalog datasets: 14
Catalog schema: 3
Layout
packages/<dataset_id>.tar: deterministic self-contained reviewed package
catalog.json: active revision heads and coverage summary… See the full description on the dataset page: https://huggingface.co/datasets/dw-indie/pad-auto-solver-reviewed.nfqa-multilingual-dataset
NFQA Multilingual Dataset
A large-scale multilingual dataset for Non-Factoid Question Answering (NFQA) classification, covering 49 languages and 8 question categories.
Dataset Statistics
Split
Examples
Train
28,653
Validation
3,539
Test
3,671
Total (Balanced)
35,863
Full Dataset (High Quality)
63,647
Dataset Composition
Languages (49 total)
Arabic (ar), Azerbaijani (az), Bulgarian (bg), Bengali (bn), Catalan (ca)… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/nfqa-multilingual-dataset.COCOPatch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
[🔗 Released Code]
[🤗 Datasets] [🤗 Checkpoints]
[📄 Tech Report] [🤗 Paper]
Figure A. PaDT pipeline.
🌟 Introduction
We are pleased to introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables multimodal large language models (MLLMs) to directly generate both textual and visual outputs.At the core of PaDT are Visual Reference Tokens (VRTs). Unlike conventional MLLMs that represent… See the full description on the dataset page: https://huggingface.co/datasets/PaDT-MLLM/COCO.ReferringImageCaptioningPatch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
[🔗 Released Code]
[🤗 Datasets] [🤗 Checkpoints]
[📄 Tech Report] [🤗 Paper]
Figure A. PaDT pipeline.
🌟 Introduction
We are pleased to introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables multimodal large language models (MLLMs) to directly generate both textual and visual outputs.At the core of PaDT are Visual Reference Tokens (VRTs). Unlike conventional MLLMs that represent… See the full description on the dataset page: https://huggingface.co/datasets/PaDT-MLLM/ReferringImageCaptioning.GSM8K_distilled_zh
Dataset
GSM8K_distilled_zh is a Chinese dataset designed for mathematical reasoning, which has been processed using MetaMath.
The question-answer pairs within this dataset have been translated from the original GSM8K dataset (available at https://github.com/openai/grade-school-math/tree/master) utilizing GPT-3.5-Turbo with few-shot prompting techniques.
This dataset comprises 7,473 training samples and 1,319 testing samples. The training samples are intended for supervised… See the full description on the dataset page: https://huggingface.co/datasets/PaddlePaddle/GSM8K_distilled_zh.webfaq-v2-bitextsWebFAQ 2.0 Bilingual Datasets (Bitexts)
Overview |
What's New in v2.0 |
Details |
Construction Method |
Structure |
Examples |
Considerations |
License |
Citation |
Contact
Note:
Note that the SIGIR Resource submission reports 104 languages, however, after re-uploading the WebFAQ 2.0 dataset, it now includes 108 languages in total.
Furthermore note that for the Bilingual Datasets, we now include all those… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq-v2-bitexts.corect-climate-feverPADBen
PADBen: Paraphrase and AI-Generated Text Detection Benchmark
📊 Dataset Overview
PADBen is a comprehensive benchmark for evaluating AI-generated text detection methods, specifically designed to test detection capabilities across various paraphrasing scenarios and attack vectors. For detailed implementation of how this dataset is generated/curated, please see https://github.com/JonathanZha47/PadBen-Paraphrase-Attack-Benchmark.
Total Dataset Size: 486,990 samples across 46… See the full description on the dataset page: https://huggingface.co/datasets/JonathanZha/PADBen.swemera-10tasks-pyconfHere’s your Markdown text, organized for clear readability:
Task Description
Instances
Instance ID
Short Title
reframe-0
Performance threshold goes to -inf when it should be zero.
pyflakes-1
Walrus operator + annotation can cause F821
sqlglot-2
MySQL dialect fails to parse PRIMARY KEY USING BTREE syntax
matchms-3
matchms fails when reading spectra where abundance is in scientific notation #809
guarddog-4
Add Mach-O magic bytes to bundled binary detector… See the full description on the dataset page: https://huggingface.co/datasets/padamenko/swemera-10tasks-pyconf.kilt-nqpad_trainThis dataset inculdes the error code of the self-refine task in the paper PaD: Program-aided Distillation Can Teach Small Models Reasoning Better than Chain-of-thought Fine-tuning.
GitHub 🔗
PAD3-Dataset-Revisi-Fixed-TRL
PAD3-Dataset-Revisi-Fixed (TRL chat format)
Conversational (TRL / SFT) dataset for age-rating classification of images.
Structure
.
├── metadata.jsonl # one TRL chat record per line
└── images/
├── Semua_Umur/000000.jpg
├── 7_/000000.jpg
├── 13_/000000.jpg
├── 15_/000000.jpg
├── 18_/000000.jpg
└── Konten_Terlarang/000000.jpg
Images are split into per-rating subfolders to stay under the 10,000-files-per-folder limit.… See the full description on the dataset page: https://huggingface.co/datasets/capstone-pad3/PAD3-Dataset-Revisi-Fixed-TRL.interview_followup_questionsPADBen-Task1
PADBen Task 1: Paraphrase Source Attribution (Binary Classification)
📋 Dataset Summary
PADBen Task 1 is a binary classification dataset for distinguishing between human-authored and LLM-generated paraphrases. This task evaluates whether AI detectors can identify the source of paraphrased text without additional context.
Key Features
Task Type: Binary text classification
Total Samples: 16,233 sentences
Train Split: 12,986 samples (80%)
Test Split: 3,247… See the full description on the dataset page: https://huggingface.co/datasets/JonathanZha/PADBen-Task1.FineCorpus-WorkoutExercise
FineCorpus-WorkoutExercise
This dataset contains structured workout exercise prompts for fine-tuning LLMs.
Structure:
conversations: Contains multi-turn dialogue pairs.
source: Indicates whether the data is from reasoning (Human) or generated by an AI model (LLM).
category: Categorizes data into Q&A, Explain, Describe, Translate.
Usage:
To use this dataset:
from datasets import load_dataset
dataset = load_dataset("padiflm/FineCorpus-WorkoutExercise"… See the full description on the dataset page: https://huggingface.co/datasets/padilfm/FineCorpus-WorkoutExercise.avvaiyar-4_kodi_padalkal
Dataset Card for Naalu Kodi Paadalgal (நாலு கோடிப் பாடல்கள்)
Summary
Naalu Kodi Paadalgal refers to a set of four famous standalone verses (Thanippaadal) attributed to the legendary poetess Avvaiyar.
The title is based on a clever wordplay. According to folklore, when challenged to compose "four crores" (Naalu Kodi) of songs in a short time, Avvaiyar composed four verses, each ending with the word "Kodi" (Crore), thus literally fulfilling the challenge of "Four-Crore… See the full description on the dataset page: https://huggingface.co/datasets/TamilThagaval/avvaiyar-4_kodi_padalkal.phyworld-data-pad_featurespaddington_en_zeroKomdigiITS-PAD2-KeywordGeneratorpad2_keywordgeneratorKomdigiITS-PAD1-Text-V2padel-rules-sft
padel-rules-sft
1,496 supervised fine-tuning examples teaching a small model padel rule fidelity:
answers accurate to the FIP regulations that import nothing from tennis or squash.
Used to train https://huggingface.co/vevag/padel-qwen3-1.7b-lora — 19.4% → 45.2% spec
adherence on a held-out set of 31 scenarios.
Format
One JSON object per line, chat format:
{"messages": [
{"role": "system", "content": "You are a helpful assistant for padel players. Answer… See the full description on the dataset page: https://huggingface.co/datasets/vevag/padel-rules-sft.ludii-instruction-answer
