datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reclorhttps://whyu.me/reclor/
@inproceedings{yu2020reclor,
author = {Yu, Weihao and Jiang, Zihang and Dong, Yanfei and Feng, Jiashi},
title = {ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning},
booktitle = {International Conference on Learning Representations (ICLR)},
month = {April},
year = {2020}
}
strategy-qafoliohttps://github.com/Yale-LILY/FOLIO
@article{han2022folio,
title={FOLIO: Natural Language Reasoning with First-Order Logic},
author = {Han, Simeng and Schoelkopf, Hailey and Zhao, Yilun and Qi, Zhenting and Riddell, Martin and Benson, Luke and Sun, Lucy and Zubova, Ekaterina and Qiao, Yujie and Burtell, Matthew and Peng, David and Fan, Jonathan and Liu, Yixin and Wong, Brian and Sailor, Malcolm and Ni, Ansong and Nan, Linyong and Kasai, Jungo and Yu, Tao and Zhang, Rui and Joty, Shafiq and… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/folio.docinsights-2026-shared-task-data
DocInsights 2026 Shared Task: DocSem
Document-grounded quantitative reasoning with evidence attribution
DocSem is the shared task of DocInsights 2026, the Workshop on Document Intelligence and Understanding co-located with EMNLP 2026 in Budapest, Hungary. The workshop theme is Beyond Plain Text: Bridging NLP and Document AI.
Workshop shared task | Source repository | Submission portal | Participant guide
Participants receive a PDF document and a paraphrased user_query. Systems… See the full description on the dataset page: https://huggingface.co/datasets/amitbcp/docinsights-2026-shared-task-data.webgym_tasks
WebGym Tasks Dataset
Dataset Description
This dataset contains web navigation tasks for training and evaluating autonomous web agents. Each task consists of a natural language instruction that describes an action to be performed on a specific website, along with evaluation criteria and metadata.
Dataset Summary
Total Training Tasks: 292,092
Total Test Tasks: 1,167
Domains: Multiple domains including Lifestyle & Leisure, Sports & Fitness, and more
Source… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/webgym_tasks.finance-tasks
Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024)
This repo contains the evaluation datasets for our paper Adapting Large Language Models via Reading Comprehension.
We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a simple method to… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/finance-tasks.osworld_v2_tasks
OSWorld V2 Task Classes
This gated dataset contains the official root-level task_*.py Python task classes for OSWorld V2.
The public GitHub repository keeps the task loader, helper utilities, and documentation. The task implementations are gated to reduce benchmark leakage and to help prevent evaluated agents from finding task answers, setup logic, or evaluator details online while executing a task.
Download from the public repository root with:
uvx --from huggingface_hub hf… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/osworld_v2_tasks.logiqa-2.0-nlihttps://github.com/csitfun/LogiQA2.0
Temporary citation:
@article{liu2020logiqa,
title={Logiqa: A challenge dataset for machine reading comprehension with logical reasoning},
author={Liu, Jian and Cui, Leyang and Liu, Hanmeng and Huang, Dandan and Wang, Yile and Zhang, Yue},
journal={arXiv preprint arXiv:2007.08124},
year={2020}
}
oscartaskyxent-tasks-gpt2race-cRace-C : additional data for race (high school/middle school) but for college level
https://github.com/mrcdata/race-c
@InProceedings{pmlr-v101-liang19a,
title={A New Multi-choice Reading Comprehension Dataset for Curriculum Learning},
author={Liang, Yichan and Li, Jianheng and Yin, Jian},
booktitle={Proceedings of The Eleventh Asian Conference on Machine Learning},
pages={742--757},
year={2019}
}
commonsense_qa_2.0https://github.com/allenai/csqa2
@article{talmor2022commonsenseqa,
title={CommonsenseQA 2.0: Exposing the limits of AI through gamification},
author={Talmor, Alon and Yoran, Ori and Bras, Ronan Le and Bhagavatula, Chandra and Goldberg, Yoav and Choi, Yejin and Berant, Jonathan},
journal={arXiv preprint arXiv:2201.05320},
year={2022}
}
sparse-reward-long-tasks
Sparse Reward Long Tasks
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/rmems/sparse-reward-long-tasks.MLS-Bench-Tasks
MLS-Bench Tasks
MLS-Bench is a benchmark for machine learning science. Where most agent benchmarks reward engineering one fixed instance — clean the data, tune the pipeline, climb a leaderboard — MLS-Bench asks the harder question: can an AI agent propose a new component, loss, optimizer, or training procedure whose gain transfers across settings, seeds, datasets, and scales?The benchmark contains 140 tasks across 12 ML research domains. Each task fixes a research scaffold… See the full description on the dataset page: https://huggingface.co/datasets/Bohan22/MLS-Bench-Tasks.Math-RL-Tasks
Ulam AI Math RL Tasks
Forty original, verifier-backed mathematical reasoning tasks packaged as ten
independent RL environments. The collection spans advanced graduate exercises,
research-style exact computation and structural generalization problems in
algebraic geometry, arithmetic geometry, combinatorics, topology, probability
and spectral analysis.
Each suite pairs a runnable rl_env/ with a preserved blind_run/ by
GPT-5.6 Sol Pro. The model name describes the evaluation actor… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/Math-RL-Tasks.c4tasky_v2SLT-Task2-Post-ASR-Speaker-Tagging
Dataset Name: Dataset for ASR Speaker-Tagging Corrections (Speaker Diarization)
Description
This dataset is pairs of erroneous ASR output and speaker tagging, which are generated from a ASR system and speaker diarization system.
Each source erroneous transcription is paired with human-annotated transcription, which has correct transcription and speaker tagging.
SEGment-wise Long-form Speech Transcription annotation (SegLST), the file format used in the CHiME challenges… See the full description on the dataset page: https://huggingface.co/datasets/GenSEC-LLM/SLT-Task2-Post-ASR-Speaker-Tagging.medicine-tasks
Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024)
This repo contains the evaluation datasets for our paper Adapting Large Language Models via Reading Comprehension.
We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a simple method to… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/medicine-tasks.leandojohttps://github.com/lean-dojo/LeanDojo
@article{yang2023leandojo,
title={{LeanDojo}: Theorem Proving with Retrieval-Augmented Language Models},
author={Yang, Kaiyu and Swope, Aidan and Gu, Alex and Chalamala, Rahul and Song, Peiyang and Yu, Shixing and Godil, Saad and Prenger, Ryan and Anandkumar, Anima},
journal={arXiv preprint arXiv:2306.15626},
year={2023}
}
DCASE2026-Task5-DevSet
DCASE 2026 Task 5 Audio-Dependent Question Answering (ADQA) Development Set
This is the official Development Set for DCASE 2026 Challenge Task 5: Audio-Dependent Question Answering (ADQA).
The ADQA task focuses on addressing "Textual Hallucination" in Large Audio-Language Models (LALMs) — where models pass audio understanding benchmarks by relying on text prompts and internal linguistic priors rather than actual audio perception. ADQA introduces a rigorous evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Harland/DCASE2026-Task5-DevSet.com2sensehttps://github.com/PlusLabNLP/Com2Sense
@inproceedings{singh-etal-2021-com2sense,
title = "{COM}2{SENSE}: A Commonsense Reasoning Benchmark with Complementary Sentences",
author = "Singh, Shikhar and
Wen, Nuan and
Hou, Yu and
Alipoormolabashi, Pegah and
Wu, Te-lin and
Ma, Xuezhe and
Peng, Nanyun",
booktitle = "Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021",
month = aug,
year = "2021",
address =… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/com2sense.sciclaimeval-shared-task
SciClaimEval Shared Task: All information is available at sciclaimeval.github.io
Evaluation scripts & examples: github.com/SciClaimEval/sciclaimeval-shared-task
More Information: paper
Version Info
Please use the latest version, v1.1.
Changes from v1.0 to v1.1
Compared with v1.0, v1.1 includes the following changes.
Removed Samples
The following 20 samples have been removed:
val_tab_1594
val_tab_0067… See the full description on the dataset page: https://huggingface.co/datasets/alabnii/sciclaimeval-shared-task.mutual@inproceedings{mutual,
title = "MuTual: A Dataset for Multi-Turn Dialogue Reasoning",
author = "Cui, Leyang and Wu, Yu and Liu, Shujie and Zhang, Yue and Zhou, Ming" ,
booktitle = "Proceedings of the 58th Conference of the Association for Computational Linguistics",
year = "2020",
publisher = "Association for Computational Linguistics",
}
ConTRoL-nlihttps://github.com/csitfun/ConTRoL-dataset
@article{Liu_Cui_Liu_Zhang_2021,
title={Natural Language Inference in Context - Investigating Contextual Reasoning over Long Texts},
volume={35},
url={https://ojs.aaai.org/index.php/AAAI/article/view/17580},
DOI={10.1609/aaai.v35i15.17580},
number={15},
journal={Proceedings of the AAAI Conference on Artificial Intelligence},
author={Liu, Hanmeng and Cui, Leyang and Liu, Jian and Zhang, Yue},
year={2021},
month={May},
pages={13388-13396}
}
spartqa-mchoicehttps://github.com/HLR/SpartQA-baselines
@inproceedings{mirzaee-etal-2021-spartqa,
title = "{SPARTQA}: A Textual Question Answering Benchmark for Spatial Reasoning",
author = "Mirzaee, Roshanak and
Rajaby Faghihi, Hossein and
Ning, Qiang and
Kordjamshidi, Parisa",
booktitle = "Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
month = jun… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/spartqa-mchoice.AppWorld-Taskssapient-synth-tasksource-reclor
sapient-synth-tasksource-reclor
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 4633
Task: synthetic anonymous instruction replacement
Generation
Rows were generated with google/gemma-4-31B-it and… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-tasksource-reclor.bigym-all-tasks-3dgs-one-success
BiGym 全任务 3DGS 厨房壳 单成功轨迹集
发布状态:40/40 技术验证通过,视觉抽检通过;仓库保持,等待上游数据/壳资产再分发条款复核。
预览视频:previews/collection-preview.mp4
核心结果
官方 BiGym 任务:40/40
唯一任务:40
保存的 reward=1 episode:40
丢弃的 reward=0 候选:61(未进入成功 Parquet)
Transition / 每相机帧数:14,806 / 14,806
三相机 H.264 文件:120
相机:head 848×480、left wrist 640×480、right wrist 640×480,20 FPS
数据布局
reach/:3 个任务
long_horizon/:3 个任务
dishwasher/:10 个任务
tabletop/:24 个任务
task-manifest.csv:任务、demo UUID、seed、帧数、reward、动作哈希… See the full description on the dataset page: https://huggingface.co/datasets/eustance/bigym-all-tasks-3dgs-one-success.scinli#SciNLI: A Corpus for Natural Language Inference on Scientific Text
https://github.com/msadat3/SciNLI
@inproceedings{sadat-caragea-2022-scinli,
title = "{S}ci{NLI}: A Corpus for Natural Language Inference on Scientific Text",
author = "Sadat, Mobashir and
Caragea, Cornelia",
booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
month = may,
year = "2022",
address = "Dublin, Ireland"… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/scinli.sie-task-evidence
SIE task evidence
The recorded inputs and model responses behind the task pages on
superlinked.com, one folder per task.
Every figure published on a task page was produced by a real recorded run against
https://api.superlinked.com. This dataset holds those recordings so that anyone
can re-derive the published numbers without an API key and without spending any
inference.
How it is used
The runnable example for each task lives in the public
superlinked/sie… See the full description on the dataset page: https://huggingface.co/datasets/superlinked/sie-task-evidence.
