datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
docinsights-2026-shared-task-data
DocInsights 2026 Shared Task: DocSem
Document-grounded quantitative reasoning with evidence attribution
DocSem is the shared task of DocInsights 2026, the Workshop on Document Intelligence and Understanding co-located with EMNLP 2026 in Budapest, Hungary. The workshop theme is Beyond Plain Text: Bridging NLP and Document AI.
Workshop shared task | Source repository | Submission portal | Participant guide
Participants receive a PDF document and a paraphrased user_query. Systems… See the full description on the dataset page: https://huggingface.co/datasets/amitbcp/docinsights-2026-shared-task-data.CoT-Reasoning-Instruct
reasoning-0.01 subset
synthetic dataset of reasoning chains for a wide variety of tasks.
we leverage data like this across multiple reasoning experiments/projects.
stay tuned for reasoning models and more data.
Thanks to Hive Digital Technologies (https://x.com/HIVEDigitalTech) for their compute support in this project and beyond.
sciclaimeval-shared-task
SciClaimEval Shared Task: All information is available at sciclaimeval.github.io
Evaluation scripts & examples: github.com/SciClaimEval/sciclaimeval-shared-task
More Information: paper
Version Info
Please use the latest version, v1.1.
Changes from v1.0 to v1.1
Compared with v1.0, v1.1 includes the following changes.
Removed Samples
The following 20 samples have been removed:
val_tab_1594
val_tab_0067… See the full description on the dataset page: https://huggingface.co/datasets/alabnii/sciclaimeval-shared-task.lmsys-chat-1m
LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
This dataset contains one million real-world conversations with 25 state-of-the-art LLMs.
It is collected from 210K unique IP addresses in the wild on the Vicuna demo and Chatbot Arena website from April to August 2023.
Each sample includes a conversation ID, model name, conversation text in OpenAI API JSON format, detected language tag, and OpenAI moderation API tag.
User consent is obtained through the "Terms of use"… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/lmsys-chat-1m.AraSeg-2026-Shared-Task-NP
Arabic Sentence Segmentation Shared Task 2026
For details about the shared task, evaluation scripts, leaderboard, and submission guidelines, visit:
https://www.araseg.aramlab.ai/
Dataset Summary
AraSeg is the first comprehensive benchmark for Arabic sentence segmentation.
The corpus is designed to support research on sentence segmentation in Modern Standard Arabic (MSA), particularly in settings where punctuation is inconsistent, missing, or noisy.
AraSeg contains… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AraSeg-2026-Shared-Task-NP.AraSeg-2026-Shared-Task-PA
Arabic Sentence Segmentation Shared Task 2026
For details about the shared task, evaluation scripts, leaderboard, and submission guidelines, visit:
https://www.araseg.aramlab.ai/
Dataset Summary
AraSeg is the first comprehensive benchmark for Arabic sentence segmentation.
The corpus is designed to support research on sentence segmentation in Modern Standard Arabic (MSA), particularly in settings where punctuation is inconsistent, missing, or noisy.
AraSeg contains… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AraSeg-2026-Shared-Task-PA.AraSeg-2026-Shared-Task-NoPnx-PA
Arabic Sentence Segmentation Shared Task 2026
For details about the shared task, evaluation scripts, leaderboard, and submission guidelines, visit:
https://www.araseg.aramlab.ai/
Dataset Summary
AraSeg is the first comprehensive benchmark for Arabic sentence segmentation.
The corpus is designed to support research on sentence segmentation in Modern Standard Arabic (MSA), particularly in settings where punctuation is inconsistent, missing, or noisy.
AraSeg contains… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AraSeg-2026-Shared-Task-NoPnx-PA.AraSeg-2026-Shared-Task-NoPnx-NP
Arabic Sentence Segmentation Shared Task 2026
For details about the shared task, evaluation scripts, leaderboard, and submission guidelines, visit:
https://www.araseg.aramlab.ai/
Dataset Summary
AraSeg is the first comprehensive benchmark for Arabic sentence segmentation.
The corpus is designed to support research on sentence segmentation in Modern Standard Arabic (MSA), particularly in settings where punctuation is inconsistent, missing, or noisy.
AraSeg contains… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AraSeg-2026-Shared-Task-NoPnx-NP.medical-qa-shared-task-v1-toy
Dataset Card for "medical-qa-shared-task-v1-toy"
More Information needed
aya_collection_telugumedical-qa-shared-task-v1-all
Dataset Card for "medical-qa-shared-task-v1-all"
More Information needed
COLING-2025-CHIPSAL
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/COLING-2025-CHIPSAL.Organic-Chemistry-VLM-PostTraining
Dataset Card for "Chemistry_text_to_image"
More Information needed
xlsum-subset
Dataset Card for "XL-Sum"
Dataset Summary
We present XLSum, a comprehensive and diverse dataset comprising 1.35 million professionally annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics. The dataset covers 45 languages ranging from low to high-resource, for many of which no public dataset is currently available. XL-Sum is highly abstractive, concise, and of high quality, as indicated by human and intrinsic evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/xlsum-subset.BAREC-Shared-Task-2025-sent
BAREC Shared Task 2025
Dataset Summary
BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset developed for the BAREC Shared Task 2025, focused on fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes.
The dataset is annotated at the sentence level. Document-level readability scores are derived by assigning each… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/BAREC-Shared-Task-2025-sent.COLING-2025-GENAI-3
🚨 RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors 🚨
🌐 Website, 🖥️ Github, 📝 Paper
RAID is the largest & most comprehensive dataset for evaluating AI-generated text detectors.
It contains over 10 million documents spanning 11 LLMs, 11 genres, 4 decoding strategies, and 12 adversarial attacks.
It is designed to be the go-to location for trustworthy third-party evaluation of both open-source and closed-source generated text detectors.
Load… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/COLING-2025-GENAI-3.MeetingBank-QA-Summarization-Instruct
Dataset Card for MeetingBank-QA-Summary
This dataset is introduced in LLMLingua-2 (Pan et al., 2024) and is designed to assess the performance of compressed meeting transcripts on downstream tasks such as question answering (QA) and summarization.
It includes 862 meeting transcripts from the test set of meeting transcripts introduced in MeetingBank (Hu et al, 2023) as the context, togeter with QA pairs and summaries that were generated by GPT-4 for each context transcripts.… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/MeetingBank-QA-Summarization-Instruct.sharedtask_nlpai4health_filtereduber_text-Vision-QA
Dataset Card for "uber_text_qa"
More Information needed
LID201_Devanagari_Script_Languages_Identificationmozilla-common-voice-spontaneous-speech-asr-shared-task
Mozilla Common Voice Spontaneous Speech ASR Shared Task
This repository combines the Mozilla Data Collective Common Voice spontaneous speech ASR shared-task
train/dev and test archives in one place.
Locales present across the combined train/dev and test packages: ady, aln, bas, bew, bxk,
cgg, el-CY, hch, kbd, kcn, koo, led, lke, lth, meh, mmc, pne, qxp, ruc,
rwm, sco, tob, top, ttj, ukv, ush.
Split package
Mozilla Data Collective dataset ID
Hub archive
Original MDC archive… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/mozilla-common-voice-spontaneous-speech-asr-shared-task.bionlp_shared_task_2009The BioNLP Shared Task 2009 was organized by GENIA Project and its corpora were curated based
on the annotations of the publicly available GENIA Event corpus and an unreleased (blind) section
of the GENIA Event corpus annotations, used for evaluation.BAREC-Shared-Task-2025-doc
BAREC Shared Task 2025
Dataset Summary
BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset developed for the BAREC Shared Task 2025, focused on fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes.
The dataset is annotated at the sentence level. Document-level readability scores are derived by assigning each… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/BAREC-Shared-Task-2025-doc.medical-qa-shared-task-v1-half
Dataset Card for "medical-qa-shared-task-v1-half"
More Information needed
SNLI-NLI
Dataset Card for SNLI
Dataset Summary
The SNLI corpus (version 1.0) is a collection of 570k human-written English sentence pairs manually labeled for balanced classification with the labels entailment, contradiction, and neutral, supporting the task of natural language inference (NLI), also known as recognizing textual entailment (RTE).
Supported Tasks and Leaderboards
Natural Language Inference (NLI), also known as Recognizing Textual Entailment (RTE), is the… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/SNLI-NLI.civil_comments_Safety
Dataset Card for "civil_comments"
Dataset Summary
The comments in this dataset come from an archive of the Civil Comments
platform, a commenting plugin for independent news sites. These public comments
were created from 2015 - 2017 and appeared on approximately 50 English-language
news sites across the world. When Civil Comments shut down in 2017, they chose
to make the public comments available in a lasting open archive to enable future
research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/civil_comments_Safety.Sujet-Vision-QA
Dataset Description 📊🔍
The Sujet-Finance-QA-Vision-100k is a comprehensive dataset containing over 100,000 question-answer pairs derived from more than 9,800 financial document images. This dataset is designed to support research and development in the field of financial document analysis and visual question answering.
Key Features:
🖼️ 9,801 unique financial document images
❓ 107,050 question-answer pairs
🇬🇧 English language
📄 Diverse financial document types… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/Sujet-Vision-QA.BAREC-Shared-Task-2026-sent
BAREC Shared Task 2026
Dataset Summary
BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset developed for the BAREC Shared Task 2026, focused on fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes.
The dataset is annotated at the sentence level. Document-level readability scores are derived by assigning each… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/BAREC-Shared-Task-2026-sent.sharedtask_nlpai4health_information_extraction_filteredCOLING-2025-FINNLP-FMD
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/COLING-2025-FINNLP-FMD.
