datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
2025-challenge-task-instancestask_data
QuantCodeEval
A benchmark for evaluating LLM coding agents on quantitative-strategy code
reproduction from finance research papers.
Status: Anonymous artifact for the 30-task benchmark.
Release mirrors
The release is mirrored at two anonymous locations:
Hugging Face Datasets — complete anonymous release:
https://huggingface.co/datasets/quantcodeeval/task_data
anonymous.4open.science — browseable mirror:
https://anonymous.4open.science/r/QuantCodeEval-Anonymous… See the full description on the dataset page: https://huggingface.co/datasets/quantcodeeval/task_data.ELYZA-tasks-100
ELYZA-tasks-100: 日本語instructionモデル評価データセット
Data Description
本データセットはinstruction-tuningを行ったモデルの評価用データセットです。詳細は リリースのnote記事 を参照してください。
特徴:
複雑な指示・タスクを含む100件の日本語データです。
役に立つAIアシスタントとして、丁寧な出力が求められます。
全てのデータに対して評価観点がアノテーションされており、評価の揺らぎを抑えることが期待されます。
具体的には以下のようなタスクを含みます。
要約を修正し、修正箇所を説明するタスク
具体的なエピソードから抽象的な教訓を述べるタスク
ユーザーの意図を汲み役に立つAIアシスタントとして振る舞うタスク
場合分けを必要とする複雑な算数のタスク
未知の言語からパターンを抽出し日本語訳する高度な推論を必要とするタスク
複数の指示を踏まえた上でyoutubeの対話を生成するタスク
架空の生き物や熟語に関する生成・大喜利などの想像力が求められるタスク… See the full description on the dataset page: https://huggingface.co/datasets/elyza/ELYZA-tasks-100.arct
The Argument Reasoning Comprehension Task: Identification and Reconstruction of Implicit Warrants
https://github.com/UKPLab/argument-reasoning-comprehension-task
@InProceedings{Habernal.et.al.2018.NAACL.ARCT,
title = {The Argument Reasoning Comprehension Task: Identification
and Reconstruction of Implicit Warrants},
author = {Habernal, Ivan and Wachsmuth, Henning and
Gurevych, Iryna and Stein, Benno},
publisher = {Association for… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/arct.rna-downstream-tasks
GB.RNA Benchmark Datasets
mRNA related tasks
Translation efficiency prediction from Chu et al.(2024) [1]
3 cell lines: Muscle, pc3, HEK
input sequence: 5'UTR
10-fold cross-validation split
mRNA expression level prediction from Chu et al.(2024) [1]
3 cell lines: Muscle, pc3, HEK
input sequence: 5'UTR
10-fold cross-validation split
Mean ribosome load prediction from Sample et al. (2019) [2]
input sequence: 5'UTR
ouput: mean ribosome load
the original data… See the full description on the dataset page: https://huggingface.co/datasets/genbio-ai/rna-downstream-tasks.jigsaw_toxicitygeneb-tasks
GENEB — Genomic Embedding Benchmark (task data)
Task-level sequence classification data for GENEB, a multi-task benchmark for DNA sequence encoders introduced in the paper: GENEB: Why Genomic Models Are Hard to Compare.
Paper: https://huggingface.co/papers/2606.04525
Source code: GitHub - darlednik/GENEB
Leaderboard: Hugging Face Space
GENEB evaluates frozen representations from 40 genomic foundation models across 100 tasks in 13 functional categories using a unified… See the full description on the dataset page: https://huggingface.co/datasets/darlednik/geneb-tasks.blog_authorship_corpussocial-chemestry-101simlexterminal-taskstask-matches2025-challenge-task-instancestraciehttps://github.com/allenai/aristo-leaderboard/tree/master/tracie/data
@inproceedings{ZRNKSR21,
author = {Ben Zhou and Kyle Richardson and Qiang Ning and Tushar Khot and Ashish Sabharwal and Dan Roth},
title = {Temporal Reasoning on Implicit Events from Distant Supervision},
booktitle = {NAACL},
year = {2021},
}
buat-task-2IELTS-writing-task-2-evaluationcounterfactually-augmented-snli@article{kaushik2020learning,
title={Learning the Difference that Makes a Difference with Counterfactually Augmented Data},
author={Kaushik, Divyansh and Hovy, Eduard and Lipton, Zachary C},
journal={International Conference on Learning Representations (ICLR)},
year={2020}
}
Original-alpha-suppression-task-boostwinowhyhttps://github.com/HKUST-KnowComp/WinoWhy
@inproceedings{zhang2020WinoWhy,
author = {Hongming Zhang and Xinran Zhao and Yangqiu Song},
title = {WinoWhy: A Deep Diagnosis of Essential Commonsense Knowledge for Answering Winograd Schema Challenge},
booktitle = {Proceedings of Annual Meeting of the Association for Computational Linguistics (ACL) 2020},
year = {2020}
}
implicit-hate-stg1https://github.com/SALT-NLP/implicit-hate
@inproceedings{elsherief-etal-2021-latent,
title = "Latent Hatred: A Benchmark for Understanding Implicit Hate Speech",
author = "ElSherief, Mai and
Ziems, Caleb and
Muchlinski, David and
Anupindi, Vaishnavi and
Seybolt, Jordyn and
De Choudhury, Munmun and
Yang, Diyi",
booktitle = "Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing",
month = nov,
year =… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/implicit-hate-stg1.I2D2code:
https://i2d2.allen.ai/
https://arxiv.org/abs/2212.09246
@inproceedings{Bhagavatula2022GenGen,
title={Generating Generics: Knowledge Induction with NeuroLogic and Self-Imitation},
author={Chandra Bhagavatula, Jena D. Hwang, Doug Downey, Ronan Le Bras, Ximing Lu, Lianhui Qin, Keisuke Sakaguchi, Swabha Swayamdipta, Peter West, Yejin Choi},
booktitle={arXiv},
year={2022}
}
Reverse-alpha-suppression-task-boostosworld_tasks_filescounterfactually-augmented-imdb@article{kaushik2020learning,
title={Learning the Difference that Makes a Difference with Counterfactually Augmented Data},
author={Kaushik, Divyansh and Hovy, Eduard and Lipton, Zachary C},
journal={International Conference on Learning Representations (ICLR)},
year={2020}
}
ate-task-ability-dataset
O*NET Task-to-Ability Mapping Dataset
A task-level mapping from 18,796 O*NET work tasks, spanning all 23 SOC major groups (economy-wide), to the 52 O*NET human abilities each task requires, with a graded importance weight per (task, ability) pair. The dataset contains 95,330 task-to-ability mappings.
It was built as the empirical foundation for the ATES (Agentic Task Exposure Score) framework, but stands alone for research on skill demand, automation exposure, and the division… See the full description on the dataset page: https://huggingface.co/datasets/ravishgupta/ate-task-ability-dataset.help-nlihttps://github.com/verypluming/HELP
@InProceedings{yanaka-EtAl:2019:starsem,
author = {Yanaka, Hitomi and Mineshima, Koji and Bekki, Daisuke and Inui, Kentaro and Sekine, Satoshi and Abzianidze, Lasha and Bos, Johan},
title = {HELP: A Dataset for Identifying Shortcomings of Neural Models in Monotonicity Reasoning},
booktitle = {Proceedings of the Eighth Joint Conference on Lexical and Computational Semantics (*SEM2019)},
year = {2019},
}
AES2-essay-scoringhttps://www.kaggle.com/competitions/learning-agency-lab-automated-essay-scoring-2/data
clef2025_checkthat_task1_subjectivity
CLEF‑2025 CheckThat! Lab Task 1: Subjectivity in News Articles
Systems are challenged to distinguish whether a sentence from a news article expresses the subjective view of the author behind it or presents an objective view on the covered topic instead.
This is a binary classification tasks in which systems have to identify whether a text sequence (a sentence or a paragraph) is subjective (SUBJ) or objective (OBJ).
The task comprises three settings:
Monolingual: train and test on… See the full description on the dataset page: https://huggingface.co/datasets/AIWizards/clef2025_checkthat_task1_subjectivity.clef2025-bioasq-task13BRadNLP2024_main_task
RadNLP 2024 main task: Document Classification for Lung Cancer Staging
📜 Paper
📚 Introduction
RadNLP 2024 is a shared task in the international conference NTCIR-18, organized by the National Institute of Informatics in Japan.
Management of lung cancer is based on the stage, and radiology reports provide various related information by describing medical images such as CT and MRI.
However, radiology reports do not always specify the stage… See the full description on the dataset page: https://huggingface.co/datasets/RadNLP/RadNLP2024_main_task.
