datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmluMMLU (hendrycks_test on huggingface) without auxiliary train. It is much lighter (7MB vs 162MB) and faster than the original implementation, in which auxiliary train is loaded (+ duplicated!) by default for all the configs in the original version, making it quite heavy.
We use this version in tasksource.
Reference to original dataset:
Measuring Massive Multitask Language Understanding - https://github.com/hendrycks/test
@article{hendryckstest2021,
title={Measuring Massive Multitask Language… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/mmlu.bigbenchBIG-Bench but it doesn't require the hellish dependencies (tensorflow, pypi-bigbench, protobuf) of the official version.
dataset = load_dataset("tasksource/bigbench",'movie_recommendation')
Code to reproduce:
https://colab.research.google.com/drive/1MKdLdF7oqrSQCeavAcsEnPdI85kD0LzU?usp=sharing
Datasets are capped to 50k examples to keep things light.
I also removed the default split when train was available also to save space, as default=train+val.
@article{srivastava2022beyond… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/bigbench.proofwriter
Dataset Card for "proofwriter"
More Information needed
TaskTrove
TaskTrove
v5.1 (current) — independent-review source retirement — moves 15 sources with majority or unanimous REJECT verdicts out of the default config and into deprecated/. Three blinded reviewers each sampled 10 tasks per source from all 50 v5.0 source-drop candidates, read the instructions and packaged tests, and issued independent KEEP or REJECT verdicts. The 15 retired sources received at least two REJECT votes. The active catalog changes from 93 sources and 1,674,033… See the full description on the dataset page: https://huggingface.co/datasets/open-thoughts/TaskTrove.nucleotide_transformer_downstream_tasks
Dataset Card for Dataset Name
The nucleotide_transformer_downstream_tasks dataset features the 18 downstream tasks presented in the Nucleotide Transformer paper. They consist of both binary and multi-class classification tasks that aim at providing a consistent genomics benchmark.
⚠️We note that we have revised and improved our benchmark during the peer-review process. The datasets featured in this repository are used up to this release. We highly encourage to move to the new… See the full description on the dataset page: https://huggingface.co/datasets/InstaDeepAI/nucleotide_transformer_downstream_tasks.lsat-lr
Dataset Card for "lsat-lr"
More Information needed
Multi-SWE-smith-tasksmulti_task_multi_modal_knowledge_retrieval_benchmark_M2KR
PreFLMR M2KR Dataset Card
Dataset details
Dataset type:
M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models.
We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.babi_nli
bAbi_nli
bAbI tasks recasted as natural language inference.
https://github.com/facebookarchive/bAbI-tasks
tasksource recasting code:
https://colab.research.google.com/drive/1J_RqDSw9iPxJSBvCJu-VRbjXnrEjKVvr?usp=sharing
@article{weston2015towards,
title={Towards ai-complete question answering: A set of prerequisite toy tasks},
author={Weston, Jason and Bordes, Antoine and Chopra, Sumit and Rush, Alexander M and Van Merri{\"e}nboer, Bart and Joulin, Armand and Mikolov, Tomas}… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/babi_nli.ropedia-xperience-10m-task-suite-artifacts
Ropedia Xperience-10M Task Suite Artifacts
This dataset repository stores small derived artifacts for the Ropedia
Xperience-10M task-suite project: metrics, predictions, manifests, reports,
figures, website JSON, public-safe Qwen3-Omni diagnostic outputs, and the
Cosmos3-Nano plus Cosmos3-Super diagnostic packages.
Project Identity
The Project identity mark is shared across the GitHub README, GitHub Pages
dashboard, Hugging Face Space, artifact dataset, model… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/ropedia-xperience-10m-task-suite-artifacts.lsat-rc
Dataset Card for "lsat-rc"
More Information needed
kilt_tasks
Dataset Card for KILT
Dataset Summary
KILT has been built from 11 datasets representing 5 types of tasks:
Fact-checking
Entity linking
Slot filling
Open domain QA
Dialog generation
All these datasets have been grounded in a single pre-processed Wikipedia dump, allowing for fairer and more consistent evaluation as well as enabling new task setups such as multitask and transfer learning with minimal effort. KILT also provides tools to analyze and understand the… See the full description on the dataset page: https://huggingface.co/datasets/facebook/kilt_tasks.lsat-ar
Dataset Card for "lsat-ar"
More Information needed
ScienceQA_text_only
Dataset Card for "scienceQA_text_only"
ScienceQA text-only examples (examples where no image was initially present, which means they should be doable with text-only models.)
@article{10.1007/s00799-022-00329-y,
author = {Saikh, Tanik and Ghosal, Tirthankar and Mittal, Amish and Ekbal, Asif and Bhattacharyya, Pushpak},
title = {ScienceQA: A Novel Resource for Question Answering on Scholarly Articles},
year = {2022},
journal = {Int. J. Digit. Libr.},
month = {sep}
}
nucleotide_transformer_downstream_tasks_revised
Dataset Card for Dataset Name
The nucleotide_transformer_downstream_tasks dataset features the 18 downstream tasks presented in the Nucleotide Transformer paper. They consist of both binary and multi-class classification tasks that aim at providing a consistent genomics benchmark.
We note that this is an updated version of this benchmark after the paper has been through peer-review. We highly encourage to move to this version in detriment of the older version.Keypoints about the… See the full description on the dataset page: https://huggingface.co/datasets/InstaDeepAI/nucleotide_transformer_downstream_tasks_revised.brainteasers
Dataset Card for "brainteasers"
More Information needed
product-photography-v1-tiny-prompts-tasks-collage-filteredlawma-tasks
Lawma legal classification tasks
This repository contains the legal classification tasks from Lawma.
These tasks were derived from the Supreme Court and Songer Court of Appeals databases.
See the project's GitHub repository for more details.
Please cite as:
@misc{dominguezolmedo2024lawmapowerspecializationlegal,
title={Lawma: The Power of Specialization for Legal Tasks},
author={Ricardo Dominguez-Olmedo and Vedant Nanda and Rediet Abebe and Stefan Bechtold and Christoph… See the full description on the dataset page: https://huggingface.co/datasets/ricdomolm/lawma-tasks.esci
Dataset Card for "esci"
ESCI product search dataset
https://github.com/amazon-science/esci-data/
Preprocessings:
-joined the two relevant files
-product_text aggregate all product text
-mapped esci_label to full name
@article{reddy2022shopping,
title={Shopping Queries Dataset: A Large-Scale {ESCI} Benchmark for Improving Product Search},
author={Chandan K. Reddy and Lluís Màrquez and Fran Valero and Nikhil Rao and Hugo Zaragoza and Sambaran Bandyopadhyay and Arnab Biswas and Anlu… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/esci.wmt20_mlqe_task1
Dataset Card for WMT20 - MultiLingual Quality Estimation (MLQE) Task1
Dataset Summary
From the homepage:
This shared task (part of WMT20) will build on its previous editions to further examine automatic methods for estimating the quality of neural machine translation output at run-time, without relying on reference translations. As in previous years, we cover estimation at various levels. Important elements introduced this year include: a new task where sentences are… See the full description on the dataset page: https://huggingface.co/datasets/wmt/wmt20_mlqe_task1.FACET-Terminal-Tasks-6k
FACET-Terminal-Tasks-6k
6,020 execution-grounded tasks for terminal agents, coding agents, and executable workflow research
🌐 FACET Project Website
📄 FACET Paper
💻 FACET-Terminal GitHub Repository
🤗 FACET-Terminal Models & Data
Dataset Overview
FACET-Terminal-Tasks-6k contains 6,020 public-release-ready Harbor tasks produced by FACET. Each task is an executable environment rather than a standalone prompt: it includes a natural-language instruction… See the full description on the dataset page: https://huggingface.co/datasets/FACET-Terminal/FACET-Terminal-Tasks-6k.tasksource-instruct-v0
Dataset Card for "tasksource-instruct-v0" (TSI)
Multi-task instruction-tuning data recasted from 485 of the tasksource datasets.
Dataset size is capped at 30k examples per task to foster task diversity.
!pip install tasksource, pandit
import tasksource, pandit
df = tasksource.list_tasks(instruct=True).sieve(id=lambda x: 'mmlu' not in x)
for tasks in df.id:
yield tasksource.load_task(task,instruct=True,max_rows=30_000,max_rows_eval=200)
https://github.com/sileod/tasksource… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/tasksource-instruct-v0.TaskMeAnything-v1-imageqa-2024
Dataset Card for TaskMeAnything-v1-imageqa-2024
TaskMeAnything-v1-imageqa-2024 benchmark dataset
🌐 Website | 📑 Paper | 🤗 Huggingface | 💻 Interface
If you like our project, please give us a star ⭐ on GitHub for latest update.
TaskMeAnything-v1-2024
TaskMeAnything-v1-imageqa-2024 is a benchmark for reflecting the current progress of MLMs by automatically finding tasks that SOTA MLMs struggle with using the TaskMeAnything Top-K queries.
This benchmark… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/TaskMeAnything-v1-imageqa-2024.ruletaker
Dataset Card for "ruletaker"
https://github.com/allenai/ruletaker
@inproceedings{ruletaker2020,
title = {Transformers as Soft Reasoners over Language},
author = {Clark, Peter and Tafjord, Oyvind and Richardson, Kyle},
booktitle = {Proceedings of the Twenty-Ninth International Joint Conference on
Artificial Intelligence, {IJCAI-20}},
publisher = {International Joint Conferences on Artificial Intelligence Organization},
editor = {Christian Bessiere}… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/ruletaker.Recursive-Task-Synthesis
Recursive Task Synthesis
This dataset contains 37,484 validated command-line task instances produced
through recursive task synthesis. Public identifiers are opaque and stable.
metadata/tasks.parquet: one searchable row per task instance.
metadata/shard_manifest.jsonl: TAR sizes and SHA256 checksums.
data/tasks-*.tar: sanitized runnable task packages.
The searchable task rows include:
instruction: contents of instruction.md.
task_toml: contents of task.toml.
solution:… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis.defeasible-nlihttps://github.com/rudinger/defeasible-nli
@inproceedings{rudinger2020thinking,
title={Thinking like a skeptic:
feasible inference in natural language},
author={Rudinger, Rachel and Shwartz, Vered and Hwang, Jena D and Bhagavatula, Chandra and Forbes, Maxwell and Le Bras, Ronan and Smith, Noah A and Choi, Yejin},
booktitle={Findings of the Association for Computational Linguistics: EMNLP 2020},
pages={4661--4675},
year={2020}
}
task903_deceptive_opinion_spam_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task903_deceptive_opinion_spam_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task903_deceptive_opinion_spam_classification.sem_eval_2010_task_8
Dataset Card for "sem_eval_2010_task_8"
Dataset Summary
The SemEval-2010 Task 8 focuses on Multi-way classification of semantic relations between pairs of nominals.
The task was designed to compare different approaches to semantic relation classification
and to provide a standard testbed for future research.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/SemEvalWorkshop/sem_eval_2010_task_8.task902_deceptive_opinion_spam_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task902_deceptive_opinion_spam_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task902_deceptive_opinion_spam_classification.Tmax-Tasks-Clean
Tmax-Tasks-Clean
New: longlongcheck (2026-09-14)
Use configuration longlongcheck, split longlongcheck, for the 431-task snapshot combining the selected Codex and Claude Code repairs with passing historical GPT-6 terminal solutions. The existing splits below retain their earlier data.
From the latest local 452-task repaired snapshot, this split holds out the requested 15 old-pass/current-fail tasks, four additional tasks without any passing current GPT-6 replay… See the full description on the dataset page: https://huggingface.co/datasets/Fzz1/Tmax-Tasks-Clean.
