shared-task
docinsights-2026-shared-task-data
DocInsights 2026 Shared Task: DocSem
Document-grounded quantitative reasoning with evidence attribution
DocSem is the shared task of DocInsights 2026, the Workshop on Document Intelligence and Understanding co-located with EMNLP 2026 in Budapest, Hungary. The workshop theme is Beyond Plain Text: Bridging NLP and Document AI.
Workshop shared task | Source repository | Submission portal | Participant guide
Participants receive a PDF document and a paraphrased user_query. Systems… See the full description on the dataset page: https://huggingface.co/datasets/amitbcp/docinsights-2026-shared-task-data.CoT-Reasoning-Instruct
reasoning-0.01 subset
synthetic dataset of reasoning chains for a wide variety of tasks.
we leverage data like this across multiple reasoning experiments/projects.
stay tuned for reasoning models and more data.
Thanks to Hive Digital Technologies (https://x.com/HIVEDigitalTech) for their compute support in this project and beyond.
sciclaimeval-shared-task
SciClaimEval Shared Task: All information is available at sciclaimeval.github.io
Evaluation scripts & examples: github.com/SciClaimEval/sciclaimeval-shared-task
More Information: paper
Version Info
Please use the latest version, v1.1.
Changes from v1.0 to v1.1
Compared with v1.0, v1.1 includes the following changes.
Removed Samples
The following 20 samples have been removed:
val_tab_1594
val_tab_0067… See the full description on the dataset page: https://huggingface.co/datasets/alabnii/sciclaimeval-shared-task.sharedtask2021
[!NOTE]
Dataset origin: https://github.com/disrpt/sharedtask2021
Introduction
The DISRPT 2021 shared task, co-located with CODI 2021 at EMNLP, introduces the second iteration of a cross-formalism shared task on discourse unit segmentation and connective detection, as well as the first iteration of a cross-formalism discourse relation classification task.
We provide training, development and test datasets from all available languages and treebanks in the RST, SDRT and PDTB… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/sharedtask2021.sharedtask2023
[!NOTE]
Dataset origin: https://github.com/disrpt/sharedtask2023
Introduction
The DISRPT 2023 shared task, to be held in conjunction with CODI 2023 and ACL 2023, introduces the third iteration of a cross-formalism shared task on discourse unit segmentation and connective detection, as well as the second iteration of a cross-formalism discourse relation classification task.
We will provide training, development, and test datasets from all available languages and treebanks in the… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/sharedtask2023.lmsys-chat-1m
LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
This dataset contains one million real-world conversations with 25 state-of-the-art LLMs.
It is collected from 210K unique IP addresses in the wild on the Vicuna demo and Chatbot Arena website from April to August 2023.
Each sample includes a conversation ID, model name, conversation text in OpenAI API JSON format, detected language tag, and OpenAI moderation API tag.
User consent is obtained through the "Terms of use"… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/lmsys-chat-1m.
