CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PaDaS-Lab /webfaq-retrievalWebFAQ Retrieval Dataset Overview | Details | Structure | Examples | Considerations | License | Citation | Contact | Acknowledgement Overview The WebFAQ Retrieval Dataset is a carefully filtered and curated subset of the broader WebFAQ Q&A Dataset.It is purpose-built for Information Retrieval (IR) tasks, such as training and evaluating dense or sparse retrieval models in multiple languages. Each of the… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq-retrieval.texttext-retrieval10M<n<100M10 likes6.3k downloads1y agoHugging Face02PaDaS-Lab /CoRECoRE: Controlled Retrieval Evaluation Dataset Motivation | Dataset Overview | Dataset Construction | Dataset Structure | Qrels Format | Evaluation | Citation | Links | Contact CoRE (Controlled Retrieval Evaluation) is a benchmark dataset designed for the rigorous evaluation of embedding compression techniques in information retrieval. 🔍 Motivation Embedding compression is essential for scaling… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/CoRE.texttext-retrieval10M<n<100M0 likes1.6k downloads11mo agoHugging Face03PaDaS-Lab /webfaqWebFAQ Q&A Dataset Overview | Details | Structure | Examples | Considerations | License | Citation | Contact | Acknowledgement Overview The WebFAQ Q&A Dataset is a broad-coverage corpus of 96 million natural question-answer (QA) pairs in 75 languages, gathered from FAQ pages on the web. It leverages structured schema.org FAQPage annotations, making it a unique resource for large-scale Question Answering… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq.textquestion-answering10M<n<100M24 likes1.4k downloads1y agoHugging Face04PaDaS-Lab /webfaq-bitextsWebFAQ Bilingual Datasets (Bitexts) Overview | Details | Structure | Examples | Considerations | License | Citation | Contact | Acknowledgement Overview The WebFAQ Bilingual Datasets (a.k.a. Bitexts) are derived from the WebFAQ Q&A Dataset, but instead of monolingual question-answer (QA) pairs, each entry here contains aligned QA pairs in two different languages. These alignments are created via… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq-bitexts.texttext-retrieval1M<n<10M3 likes632 downloads2y agoHugging Face05nyu-dice-lab /lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v1.0 Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v1.0 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private.tabular100K<n<1M0 likes515 downloads2y agoHugging Face06PaDT-MLLM /RefCOCOPatch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs [🔗 Released Code] [🤗 Datasets] [🤗 Checkpoints] [📄 Tech Report] [🤗 Paper] Figure A. PaDT pipeline. 🌟 Introduction We are pleased to introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables multimodal large language models (MLLMs) to directly generate both textual and visual outputs.At the core of PaDT are Visual Reference Tokens (VRTs). Unlike conventional MLLMs that represent… See the full description on the dataset page: https://huggingface.co/datasets/PaDT-MLLM/RefCOCO.textobject-detection100K<n<1M4 likes329 downloads1y agoHugging Face07nyu-dice-lab /lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v3.0 Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v3.0 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private.tabular100K<n<1M0 likes302 downloads2y agoHugging Face08dw-indie /pad-auto-solver-reviewed PAD Reviewed Dataset Canonical reviewed PAD board/orb artifacts for dw-indie/pad-auto-solver-reviewed. This repository contains immutable reviewed package revisions and does not contain raw captures, training runs, checkpoints, or model binaries. Packages exported: 28 Active catalog datasets: 14 Catalog schema: 3 Layout packages/<dataset_id>.tar: deterministic self-contained reviewed package catalog.json: active revision heads and coverage summary… See the full description on the dataset page: https://huggingface.co/datasets/dw-indie/pad-auto-solver-reviewed.tabularimage-classification10K<n<100K1 likes205 downloads16d agoHugging Face09PaDaS-Lab /nfqa-multilingual-dataset NFQA Multilingual Dataset A large-scale multilingual dataset for Non-Factoid Question Answering (NFQA) classification, covering 49 languages and 8 question categories. Dataset Statistics Split Examples Train 28,653 Validation 3,539 Test 3,671 Total (Balanced) 35,863 Full Dataset (High Quality) 63,647 Dataset Composition Languages (49 total) Arabic (ar), Azerbaijani (az), Bulgarian (bg), Bengali (bn), Catalan (ca)… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/nfqa-multilingual-dataset.tabulartext-classification10K<n<100K1 likes165 downloads6mo agoHugging Face10PaDT-MLLM /COCOPatch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs [🔗 Released Code] [🤗 Datasets] [🤗 Checkpoints] [📄 Tech Report] [🤗 Paper] Figure A. PaDT pipeline. 🌟 Introduction We are pleased to introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables multimodal large language models (MLLMs) to directly generate both textual and visual outputs.At the core of PaDT are Visual Reference Tokens (VRTs). Unlike conventional MLLMs that represent… See the full description on the dataset page: https://huggingface.co/datasets/PaDT-MLLM/COCO.textobject-detection100K<n<1M1 likes148 downloads1y agoHugging Face11PaDT-MLLM /ReferringImageCaptioningPatch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs [🔗 Released Code] [🤗 Datasets] [🤗 Checkpoints] [📄 Tech Report] [🤗 Paper] Figure A. PaDT pipeline. 🌟 Introduction We are pleased to introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables multimodal large language models (MLLMs) to directly generate both textual and visual outputs.At the core of PaDT are Visual Reference Tokens (VRTs). Unlike conventional MLLMs that represent… See the full description on the dataset page: https://huggingface.co/datasets/PaDT-MLLM/ReferringImageCaptioning.textimage-to-text100K<n<1M3 likes133 downloads1y agoHugging Face12PaddlePaddle /GSM8K_distilled_zh Dataset GSM8K_distilled_zh is a Chinese dataset designed for mathematical reasoning, which has been processed using MetaMath. The question-answer pairs within this dataset have been translated from the original GSM8K dataset (available at https://github.com/openai/grade-school-math/tree/master) utilizing GPT-3.5-Turbo with few-shot prompting techniques. This dataset comprises 7,473 training samples and 1,319 testing samples. The training samples are intended for supervised… See the full description on the dataset page: https://huggingface.co/datasets/PaddlePaddle/GSM8K_distilled_zh.text1K<n<10K2 likes114 downloads2y agoHugging Face13PaDaS-Lab /webfaq-v2-bitextsWebFAQ 2.0 Bilingual Datasets (Bitexts) Overview | What's New in v2.0 | Details | Construction Method | Structure | Examples | Considerations | License | Citation | Contact Note: Note that the SIGIR Resource submission reports 104 languages, however, after re-uploading the WebFAQ 2.0 dataset, it now includes 108 languages in total. Furthermore note that for the Bilingual Datasets, we now include all those… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq-v2-bitexts.texttext-retrieval10M<n<100M1 likes92 downloads7mo agoHugging Face14PaDaS-Lab /corect-climate-fevertexttext-retrieval1M<n<10M0 likes76 downloads5mo agoHugging Face15JonathanZha /PADBen PADBen: Paraphrase and AI-Generated Text Detection Benchmark 📊 Dataset Overview PADBen is a comprehensive benchmark for evaluating AI-generated text detection methods, specifically designed to test detection capabilities across various paraphrasing scenarios and attack vectors. For detailed implementation of how this dataset is generated/curated, please see https://github.com/JonathanZha47/PadBen-Paraphrase-Attack-Benchmark. Total Dataset Size: 486,990 samples across 46… See the full description on the dataset page: https://huggingface.co/datasets/JonathanZha/PADBen.tabulartext-classification100K<n<1M0 likes47 downloads11mo agoHugging Face16padamenko /swemera-10tasks-pyconfHere’s your Markdown text, organized for clear readability: Task Description Instances Instance ID Short Title reframe-0 Performance threshold goes to -inf when it should be zero. pyflakes-1 Walrus operator + annotation can cause F821 sqlglot-2 MySQL dialect fails to parse PRIMARY KEY USING BTREE syntax matchms-3 matchms fails when reading spectra where abundance is in scientific notation #809 guarddog-4 Add Mach-O magic bytes to bundled binary detector… See the full description on the dataset page: https://huggingface.co/datasets/padamenko/swemera-10tasks-pyconf.tabularn<1K1 likes42 downloads1y agoHugging Face17PaDaS-Lab /kilt-nqtexttext-retrieval1K<n<10K0 likes38 downloads5mo agoHugging Face18xuekai /pad_trainThis dataset inculdes the error code of the self-refine task in the paper PaD: Program-aided Distillation Can Teach Small Models Reasoning Better than Chain-of-thought Fine-tuning. GitHub 🔗 text100K<n<1M0 likes33 downloads3y agoHugging Face19capstone-pad3 /PAD3-Dataset-Revisi-Fixed-TRL PAD3-Dataset-Revisi-Fixed (TRL chat format) Conversational (TRL / SFT) dataset for age-rating classification of images. Structure . ├── metadata.jsonl # one TRL chat record per line └── images/ ├── Semua_Umur/000000.jpg ├── 7_/000000.jpg ├── 13_/000000.jpg ├── 15_/000000.jpg ├── 18_/000000.jpg └── Konten_Terlarang/000000.jpg Images are split into per-rating subfolders to stay under the 10,000-files-per-folder limit.… See the full description on the dataset page: https://huggingface.co/datasets/capstone-pad3/PAD3-Dataset-Revisi-Fixed-TRL.imageimage-text-to-text10K<n<100K0 likes18 downloads3mo agoHugging Face20paddy4544 /interview_followup_questionstext1M<n<10M2 likes16 downloads2y agoHugging Face21JonathanZha /PADBen-Task1 PADBen Task 1: Paraphrase Source Attribution (Binary Classification) 📋 Dataset Summary PADBen Task 1 is a binary classification dataset for distinguishing between human-authored and LLM-generated paraphrases. This task evaluates whether AI detectors can identify the source of paraphrased text without additional context. Key Features Task Type: Binary text classification Total Samples: 16,233 sentences Train Split: 12,986 samples (80%) Test Split: 3,247… See the full description on the dataset page: https://huggingface.co/datasets/JonathanZha/PADBen-Task1.tabulartext-classification10K<n<100K0 likes16 downloads1y agoHugging Face22padilfm /FineCorpus-WorkoutExercise FineCorpus-WorkoutExercise This dataset contains structured workout exercise prompts for fine-tuning LLMs. Structure: conversations: Contains multi-turn dialogue pairs. source: Indicates whether the data is from reasoning (Human) or generated by an AI model (LLM). category: Categorizes data into Q&A, Explain, Describe, Translate. Usage: To use this dataset: from datasets import load_dataset dataset = load_dataset("padiflm/FineCorpus-WorkoutExercise"… See the full description on the dataset page: https://huggingface.co/datasets/padilfm/FineCorpus-WorkoutExercise.texttext-generationn<1K0 likes14 downloads2y agoHugging Face23TamilThagaval /avvaiyar-4_kodi_padalkal Dataset Card for Naalu Kodi Paadalgal (நாலு கோடிப் பாடல்கள்) Summary Naalu Kodi Paadalgal refers to a set of four famous standalone verses (Thanippaadal) attributed to the legendary poetess Avvaiyar. The title is based on a clever wordplay. According to folklore, when challenged to compose "four crores" (Naalu Kodi) of songs in a short time, Avvaiyar composed four verses, each ending with the word "Kodi" (Crore), thus literally fulfilling the challenge of "Four-Crore… See the full description on the dataset page: https://huggingface.co/datasets/TamilThagaval/avvaiyar-4_kodi_padalkal.texttext-generationn<1K0 likes14 downloads6mo agoHugging Face24huiliu123 /phyworld-data-pad_featurestabularn<1K0 likes13 downloads5mo agoHugging Face25sibozhu /paddington_en_zerotextn<1K0 likes11 downloads2y agoHugging Face26aitf-komdigi /KomdigiITS-PAD2-KeywordGenerator10K<n<100K0 likes9 downloads3mo agoHugging Face27thalytavius /pad2_keywordgeneratortext1K<n<10K0 likes8 downloads5mo agoHugging Face28aitf-komdigi /KomdigiITS-PAD1-Text-V2text10K<n<100K0 likes8 downloads3mo agoHugging Face29vevag /padel-rules-sft padel-rules-sft 1,496 supervised fine-tuning examples teaching a small model padel rule fidelity: answers accurate to the FIP regulations that import nothing from tennis or squash. Used to train https://huggingface.co/vevag/padel-qwen3-1.7b-lora — 19.4% → 45.2% spec adherence on a held-out set of 31 scenarios. Format One JSON object per line, chat format: {"messages": [ {"role": "system", "content": "You are a helpful assistant for padel players. Answer… See the full description on the dataset page: https://huggingface.co/datasets/vevag/padel-rules-sft.textquestion-answering1K<n<10K0 likes7 downloads1mo agoHugging Face30Padlex /ludii-instruction-answertext1K<n<10K0 likes6 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.