CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amitbcp /docinsights-2026-shared-task-data DocInsights 2026 Shared Task: DocSem Document-grounded quantitative reasoning with evidence attribution DocSem is the shared task of DocInsights 2026, the Workshop on Document Intelligence and Understanding co-located with EMNLP 2026 in Budapest, Hungary. The workshop theme is Beyond Plain Text: Bridging NLP and Document AI. Workshop shared task | Source repository | Submission portal | Participant guide Participants receive a PDF document and a paraphrased user_query. Systems… See the full description on the dataset page: https://huggingface.co/datasets/amitbcp/docinsights-2026-shared-task-data.documentquestion-answering1K<n<10K0 likes4.9k downloads19d agoHugging Face021-800-SHARED-TASKS /CoT-Reasoning-Instruct reasoning-0.01 subset synthetic dataset of reasoning chains for a wide variety of tasks. we leverage data like this across multiple reasoning experiments/projects. stay tuned for reasoning models and more data. Thanks to Hive Digital Technologies (https://x.com/HIVEDigitalTech) for their compute support in this project and beyond. text10K<n<100K2 likes551 downloads2y agoHugging Face03alabnii /sciclaimeval-shared-task SciClaimEval Shared Task: All information is available at sciclaimeval.github.io Evaluation scripts & examples: github.com/SciClaimEval/sciclaimeval-shared-task More Information: paper Version Info Please use the latest version, v1.1. Changes from v1.0 to v1.1 Compared with v1.0, v1.1 includes the following changes. Removed Samples The following 20 samples have been removed: val_tab_1594 val_tab_0067… See the full description on the dataset page: https://huggingface.co/datasets/alabnii/sciclaimeval-shared-task.imagetext-classification1K<n<10K4 likes442 downloads1mo agoHugging Face041-800-SHARED-TASKS /lmsys-chat-1m LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset This dataset contains one million real-world conversations with 25 state-of-the-art LLMs. It is collected from 210K unique IP addresses in the wild on the Vicuna demo and Chatbot Arena website from April to August 2023. Each sample includes a conversation ID, model name, conversation text in OpenAI API JSON format, detected language tag, and OpenAI moderation API tag. User consent is obtained through the "Terms of use"… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/lmsys-chat-1m.text1M<n<10M1 likes174 downloads2y agoHugging Face05MBZUAI /AraSeg-2026-Shared-Task-NP Arabic Sentence Segmentation Shared Task 2026 For details about the shared task, evaluation scripts, leaderboard, and submission guidelines, visit: https://www.araseg.aramlab.ai/ Dataset Summary AraSeg is the first comprehensive benchmark for Arabic sentence segmentation. The corpus is designed to support research on sentence segmentation in Modern Standard Arabic (MSA), particularly in settings where punctuation is inconsistent, missing, or noisy. AraSeg contains… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AraSeg-2026-Shared-Task-NP.texttoken-classificationn<1K1 likes170 downloads4mo agoHugging Face06MBZUAI /AraSeg-2026-Shared-Task-PA Arabic Sentence Segmentation Shared Task 2026 For details about the shared task, evaluation scripts, leaderboard, and submission guidelines, visit: https://www.araseg.aramlab.ai/ Dataset Summary AraSeg is the first comprehensive benchmark for Arabic sentence segmentation. The corpus is designed to support research on sentence segmentation in Modern Standard Arabic (MSA), particularly in settings where punctuation is inconsistent, missing, or noisy. AraSeg contains… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AraSeg-2026-Shared-Task-PA.texttoken-classificationn<1K1 likes153 downloads4mo agoHugging Face07MBZUAI /AraSeg-2026-Shared-Task-NoPnx-PA Arabic Sentence Segmentation Shared Task 2026 For details about the shared task, evaluation scripts, leaderboard, and submission guidelines, visit: https://www.araseg.aramlab.ai/ Dataset Summary AraSeg is the first comprehensive benchmark for Arabic sentence segmentation. The corpus is designed to support research on sentence segmentation in Modern Standard Arabic (MSA), particularly in settings where punctuation is inconsistent, missing, or noisy. AraSeg contains… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AraSeg-2026-Shared-Task-NoPnx-PA.texttoken-classificationn<1K2 likes152 downloads4mo agoHugging Face08MBZUAI /AraSeg-2026-Shared-Task-NoPnx-NP Arabic Sentence Segmentation Shared Task 2026 For details about the shared task, evaluation scripts, leaderboard, and submission guidelines, visit: https://www.araseg.aramlab.ai/ Dataset Summary AraSeg is the first comprehensive benchmark for Arabic sentence segmentation. The corpus is designed to support research on sentence segmentation in Modern Standard Arabic (MSA), particularly in settings where punctuation is inconsistent, missing, or noisy. AraSeg contains… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AraSeg-2026-Shared-Task-NoPnx-NP.texttoken-classificationn<1K1 likes149 downloads4mo agoHugging Face09lavita /medical-qa-shared-task-v1-toy Dataset Card for "medical-qa-shared-task-v1-toy" More Information needed tabularn<1K41 likes146 downloads3y agoHugging Face101-800-SHARED-TASKS /aya_collection_telugutabular1M<n<10M0 likes108 downloads2y agoHugging Face11lavita /medical-qa-shared-task-v1-all Dataset Card for "medical-qa-shared-task-v1-all" More Information needed tabular10K<n<100K4 likes94 downloads3y agoHugging Face121-800-SHARED-TASKS /COLING-2025-CHIPSAL Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/COLING-2025-CHIPSAL.tabular100K<n<1M2 likes87 downloads2y agoHugging Face131-800-SHARED-TASKS /Organic-Chemistry-VLM-PostTraining Dataset Card for "Chemistry_text_to_image" More Information needed image100K<n<1M6 likes67 downloads2y agoHugging Face141-800-SHARED-TASKS /xlsum-subset Dataset Card for "XL-Sum" Dataset Summary We present XLSum, a comprehensive and diverse dataset comprising 1.35 million professionally annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics. The dataset covers 45 languages ranging from low to high-resource, for many of which no public dataset is currently available. XL-Sum is highly abstractive, concise, and of high quality, as indicated by human and intrinsic evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/xlsum-subset.textsummarizationn<1K0 likes59 downloads2y agoHugging Face15CAMeL-Lab /BAREC-Shared-Task-2025-sent BAREC Shared Task 2025 Dataset Summary BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset developed for the BAREC Shared Task 2025, focused on fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes. The dataset is annotated at the sentence level. Document-level readability scores are derived by assigning each… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/BAREC-Shared-Task-2025-sent.tabulartext-classification10K<n<100K2 likes51 downloads1y agoHugging Face161-800-SHARED-TASKS /COLING-2025-GENAI-3 🚨 RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors 🚨 🌐 Website, 🖥️ Github, 📝 Paper RAID is the largest & most comprehensive dataset for evaluating AI-generated text detectors. It contains over 10 million documents spanning 11 LLMs, 11 genres, 4 decoding strategies, and 12 adversarial attacks. It is designed to be the go-to location for trustworthy third-party evaluation of both open-source and closed-source generated text detectors. Load… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/COLING-2025-GENAI-3.texttext-classification1M<n<10M0 likes50 downloads2y agoHugging Face171-800-SHARED-TASKS /MeetingBank-QA-Summarization-Instruct Dataset Card for MeetingBank-QA-Summary This dataset is introduced in LLMLingua-2 (Pan et al., 2024) and is designed to assess the performance of compressed meeting transcripts on downstream tasks such as question answering (QA) and summarization. It includes 862 meeting transcripts from the test set of meeting transcripts introduced in MeetingBank (Hu et al, 2023) as the context, togeter with QA pairs and summaries that were generated by GPT-4 for each context transcripts.… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/MeetingBank-QA-Summarization-Instruct.textquestion-answeringn<1K1 likes49 downloads2y agoHugging Face18vinaybabu /sharedtask_nlpai4health_filteredtext10K<n<100K0 likes47 downloads1y agoHugging Face191-800-SHARED-TASKS /uber_text-Vision-QA Dataset Card for "uber_text_qa" More Information needed image1K<n<10K0 likes45 downloads2y agoHugging Face201-800-SHARED-TASKS /LID201_Devanagari_Script_Languages_Identificationtext1M<n<10M0 likes42 downloads2y agoHugging Face21Peacockery /mozilla-common-voice-spontaneous-speech-asr-shared-task Mozilla Common Voice Spontaneous Speech ASR Shared Task This repository combines the Mozilla Data Collective Common Voice spontaneous speech ASR shared-task train/dev and test archives in one place. Locales present across the combined train/dev and test packages: ady, aln, bas, bew, bxk, cgg, el-CY, hch, kbd, kcn, koo, led, lke, lth, meh, mmc, pne, qxp, ruc, rwm, sco, tob, top, ttj, ukv, ush. Split package Mozilla Data Collective dataset ID Hub archive Original MDC archive… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/mozilla-common-voice-spontaneous-speech-asr-shared-task.textautomatic-speech-recognition10K<n<100K0 likes39 downloads3mo agoHugging Face22bigbio /bionlp_shared_task_2009The BioNLP Shared Task 2009 was organized by GENIA Project and its corpora were curated based on the annotations of the publicly available GENIA Event corpus and an unreleased (blind) section of the GENIA Event corpus annotations, used for evaluation.text1K<n<10K2 likes37 downloads4y agoHugging Face23CAMeL-Lab /BAREC-Shared-Task-2025-doc BAREC Shared Task 2025 Dataset Summary BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset developed for the BAREC Shared Task 2025, focused on fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes. The dataset is annotated at the sentence level. Document-level readability scores are derived by assigning each… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/BAREC-Shared-Task-2025-doc.tabulartext-classification1K<n<10K2 likes35 downloads1y agoHugging Face24lavita /medical-qa-shared-task-v1-half Dataset Card for "medical-qa-shared-task-v1-half" More Information needed tabular1K<n<10K1 likes32 downloads3y agoHugging Face251-800-SHARED-TASKS /SNLI-NLI Dataset Card for SNLI Dataset Summary The SNLI corpus (version 1.0) is a collection of 570k human-written English sentence pairs manually labeled for balanced classification with the labels entailment, contradiction, and neutral, supporting the task of natural language inference (NLI), also known as recognizing textual entailment (RTE). Supported Tasks and Leaderboards Natural Language Inference (NLI), also known as Recognizing Textual Entailment (RTE), is the… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/SNLI-NLI.texttext-classification100K<n<1M0 likes31 downloads2y agoHugging Face261-800-SHARED-TASKS /civil_comments_Safety Dataset Card for "civil_comments" Dataset Summary The comments in this dataset come from an archive of the Civil Comments platform, a commenting plugin for independent news sites. These public comments were created from 2015 - 2017 and appeared on approximately 50 English-language news sites across the world. When Civil Comments shut down in 2017, they chose to make the public comments available in a lasting open archive to enable future research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/civil_comments_Safety.tabulartext-classification1M<n<10M0 likes31 downloads2y agoHugging Face271-800-SHARED-TASKS /Sujet-Vision-QA Dataset Description 📊🔍 The Sujet-Finance-QA-Vision-100k is a comprehensive dataset containing over 100,000 question-answer pairs derived from more than 9,800 financial document images. This dataset is designed to support research and development in the field of financial document analysis and visual question answering. Key Features: 🖼️ 9,801 unique financial document images ❓ 107,050 question-answer pairs 🇬🇧 English language 📄 Diverse financial document types… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/Sujet-Vision-QA.imagequestion-answering1K<n<10K0 likes31 downloads2y agoHugging Face28CAMeL-Lab /BAREC-Shared-Task-2026-sent BAREC Shared Task 2026 Dataset Summary BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset developed for the BAREC Shared Task 2026, focused on fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes. The dataset is annotated at the sentence level. Document-level readability scores are derived by assigning each… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/BAREC-Shared-Task-2026-sent.tabulartext-classification10K<n<100K1 likes30 downloads4mo agoHugging Face29vinaybabu /sharedtask_nlpai4health_information_extraction_filteredtext10K<n<100K0 likes29 downloads1y agoHugging Face301-800-SHARED-TASKS /COLING-2025-FINNLP-FMD Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/COLING-2025-FINNLP-FMD.text1K<n<10K1 likes28 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.