datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
docinsights-2026-shared-task-data
DocInsights 2026 Shared Task: DocSem
Document-grounded quantitative reasoning with evidence attribution
DocSem is the shared task of DocInsights 2026, the Workshop on Document Intelligence and Understanding co-located with EMNLP 2026 in Budapest, Hungary. The workshop theme is Beyond Plain Text: Bridging NLP and Document AI.
Workshop shared task | Source repository | Submission portal | Participant guide
Participants receive a PDF document and a paraphrased user_query. Systems… See the full description on the dataset page: https://huggingface.co/datasets/amitbcp/docinsights-2026-shared-task-data.sciclaimeval-shared-task
SciClaimEval Shared Task: All information is available at sciclaimeval.github.io
Evaluation scripts & examples: github.com/SciClaimEval/sciclaimeval-shared-task
More Information: paper
Version Info
Please use the latest version, v1.1.
Changes from v1.0 to v1.1
Compared with v1.0, v1.1 includes the following changes.
Removed Samples
The following 20 samples have been removed:
val_tab_1594
val_tab_0067… See the full description on the dataset page: https://huggingface.co/datasets/alabnii/sciclaimeval-shared-task.TSAR2025_SharedTask_RCTS_Test-Data
Citation
@inproceedings{alva-manchego-etal-2025-findings,
title = "Findings of the {TSAR} 2025 Shared Task on Readability-Controlled Text Simplification",
author = "Alva-Manchego, Fernando and Stodden, Regina and Imperial, Joseph Marvin and Barayan, Abdullah and North, Kai and Tayyar Madabushi, Harish",
editor = "Shardlow, Matthew and Alva-Manchego, Fernando and North, Kai and Stodden, Regina and Saggion, Horacio and Khallaf, Nouran and Hayakawa, Akio"… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/TSAR2025_SharedTask_RCTS_Test-Data.sciclaimeval-shared-task
SciClaimEval Shared Task: All information is available at sciclaimeval.github.io
Evaluation scripts & examples: github.com/SciClaimEval/sciclaimeval-shared-task
More Information: paper
Note on License Information
The dataset is licensed under CC BY 4.0; however, individual samples may have their own licenses.
MNLI-NLI
Glue MNLI
This dataset is a port of the official mnli dataset on the Hub.
It contains the matched version.
Note that the premise and hypothesis columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
sciclaimeval-shared-task-test
Grant Access: This dataset is available only to NTCIR participants.
SciClaimEval Shared Task: All information is available at sciclaimeval.github.io
Version Info
Please use the latest version, v1.1.
Changes from v1.0 to v1.1
Compared with v1.0, v1.1 removes the following 11 samples:
test_fig_0119
test_fig_0421
test_tab_0189
test_tab_0240
test_tab_0256
test_fig_0172
test_fig_0173
test_fig_0430
test_tab_0137
test_fig_0111
test_tab_0052… See the full description on the dataset page: https://huggingface.co/datasets/alabnii/sciclaimeval-shared-task-test.BPCC_TeluguTSAR2025_SharedTask_RCTS_Trial-Data
Citation
@inproceedings{alva-manchego-etal-2025-findings,
title = "Findings of the {TSAR} 2025 Shared Task on Readability-Controlled Text Simplification",
author = "Alva-Manchego, Fernando and Stodden, Regina and Imperial, Joseph Marvin and Barayan, Abdullah and North, Kai and Tayyar Madabushi, Harish",
editor = "Shardlow, Matthew and Alva-Manchego, Fernando and North, Kai and Stodden, Regina and Saggion, Horacio and Khallaf, Nouran and Hayakawa, Akio"… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/TSAR2025_SharedTask_RCTS_Trial-Data.allenai-multipref
MultiPref - a multi-annotated and multi-aspect human preference dataset
[Paper: coming soon!]
Dataset Summary
The MultiPref dataset (version 1.0) is a rich collection of 10k human preferences. It is:
Multi-annotated: each instance is annotated multiple times—twice by normal crowdworkers and twice by domain-experts— resulting in around 40k annotations.
Multi-aspect: aside from their Overall preference, annotators choose their preferred response on a five-point… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/allenai-multipref.
