datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task1577_amazon_reviews_multi_japanese_language_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1577_amazon_reviews_multi_japanese_language_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1577_amazon_reviews_multi_japanese_language_classification.task638_multi_woz_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task638_multi_woz_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task638_multi_woz_classification.multi-strategy-algorithmic-tasks
Multi-Strategy Algorithmic Tasks
A synthetic benchmark of parseable algorithmic problems with multiple valid
solution strategies for each task. Each example contains a problem,a strategy-specific
solution trace, and the strategy used to generate that trace.
The benchmark accompanies
Uncovering Latent Reasoning Strategies in Language Models,
which studies the problem of recovering mixtures of strategies implicitly represented in language models.
The benchmark provides a… See the full description on the dataset page: https://huggingface.co/datasets/awni00/multi-strategy-algorithmic-tasks.Kazakh_Multi-Task_corpus
Kazakh Multi-task Corpus
A multi-task NLP dataset in the Kazakh language, covering seven distinct language tasks - from instruction-following and question answering to translation, sentiment analysis, and grammar exercises. Designed to support the development of Kazakh-language models, benchmarks, and linguistic research.
Dataset Summary
Kazakh is a Turkic language spoken by over 13 million people, yet it remains significantly underrepresented in NLP research and… See the full description on the dataset page: https://huggingface.co/datasets/mangi-llm/Kazakh_Multi-Task_corpus.configurable-system-prompt-multitask
Configurable System Prompt Multi-task Dataset 🛞
We release the synthetic dataset for the multi-task experiments from the paper "Configurable Safety Tuning of Language Models with Synthetic Preference Data", https://huggingface.co/papers/2404.00495. This dataset has two sources for the examples:
Self-critique on a safety task from Harmful Behaviours, using the SOLAR-Instruct model. It employs two system prompts to learn the different behaviors:
You are a helpful yet harmless… See the full description on the dataset page: https://huggingface.co/datasets/vicgalle/configurable-system-prompt-multitask.multilingual-multitask-refusal
Multilingual Multitask Refusal
A multilingual, multitask prompt dataset for analysing refusal behaviour across language, task wrapper, and harmful / harmless labels.
English seed spans and labels come from the previous Multitask Multilingual Refusal dataset. Non-English content is produced with Google Sheets GOOGLETRANSLATE. Task instructions are language-localized via templates_localized.json.
Rows
211,320
English seeds
1,761
Languages
15
Tasks
8
Product
1… See the full description on the dataset page: https://huggingface.co/datasets/kaustubhg73/multilingual-multitask-refusal.Telugu-MultiTask-Instruct-77K
Telugu MultiTask Instruct 77K — Adaption AutoScientist Challenge Dataset
Powered by Adaptive Data — Adaption Labs
Dataset Description
A large-scale, multi-task Telugu instruction-tuning dataset combining 77,653 rows from 7 open-source Telugu NLP collections. Covers diverse tasks including news summarization, QA, creative writing, translation, and general instruction following — all processed through the Adaption Labs AutoScientist platform for quality… See the full description on the dataset page: https://huggingface.co/datasets/narendarcodes/Telugu-MultiTask-Instruct-77K.sec-extraction-multitask-v4
SEC Extraction Multitask v4
Instruction-tuning dataset for fine-tuning a small language model (e.g. Gemma 4 E2B) to extract structured data from SEC filings across three verticals:
Exhibit 10 (contracts) — financial terms from executive employment, credit agreements, indemnification, licensing, and similar filings
DEF 14A (proxy statements) — executive compensation, governance items, say-on-pay
MD&A (10-K / 10-Q Management's Discussion & Analysis) — operating metrics, segment… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-extraction-multitask-v4.thai-multitask-starter
Thai Multitask 9.6K
ชุดข้อมูลตั้งต้นสำหรับ instruction tuning ภาษาไทย ครอบคลุมงานสนทนา ถาม–ตอบ สรุป
แปล จำแนกข้อความ ตรวจแก้ภาษา คณิตศาสตร์ และ structured output
ข้อมูลทุกแถวสร้างขึ้นใหม่ด้วยกฎแบบ deterministic ไม่มีการคัดลอกจากเว็บไซต์หรือ
ข้อมูลส่วนบุคคลจริง เหมาะสำหรับทดลอง supervised fine-tuning และทดสอบ pipeline
แต่ควรเพิ่มข้อมูลที่มนุษย์ตรวจทานและข้อมูลภาษาธรรมชาติก่อนใช้กับระบบจริง
จำนวนข้อมูลทั้งหมด 9,599 ตัวอย่าง: train 8,639, validation 480 และ test 480… See the full description on the dataset page: https://huggingface.co/datasets/Phettae/thai-multitask-starter.task1575_amazon_reviews_multi_sentiment_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1575_amazon_reviews_multi_sentiment_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1575_amazon_reviews_multi_sentiment_classification.task639_multi_woz_user_utterance_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task639_multi_woz_user_utterance_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task639_multi_woz_user_utterance_generation.scientific-multitask-instructions
Scientific Multitask Instructions
A multi-task scientific instruction-following dataset created for
supervised fine-tuning and preference-optimization experiments.
Dataset summary
The dataset contains 1,576 conversational scientific examples across
eight task types.
Split
Examples
Train
1,260
Validation
158
Test
158
Total
1,576
Task distribution
Task
Examples
Scientific question answering
256
Summarization
220… See the full description on the dataset page: https://huggingface.co/datasets/Miladsaeedi70/scientific-multitask-instructions.task1576_amazon_reviews_multi_english_language_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1576_amazon_reviews_multi_english_language_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1576_amazon_reviews_multi_english_language_classification.sandman-dream_multitask_train
Sandman dream multitask v1 — train split
The first version of the train split used to train
sandman-gemma3-1b-multitask.
Superseded by v2,
a smaller, more curated set built on
DreamBank
rather than this one's broader source. Kept here for reference.
multi-task-instructionsandman-dream_multitask_test
Sandman dream multitask v1 — test split
The first version of the test split used to train
sandman-gemma3-1b-multitask.
Superseded by v2,
a smaller, more curated set built on
DreamBank
rather than this one's broader source. Kept here for reference.
task1574_amazon_reviews_multi_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1574_amazon_reviews_multi_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1574_amazon_reviews_multi_language_identification.sandman-dream_multitask_val
Sandman dream multitask v1 — val split
The first version of the val split used to train
sandman-gemma3-1b-multitask.
Superseded by v2,
a smaller, more curated set built on
DreamBank
rather than this one's broader source. Kept here for reference.
sandman-dream_multitask_v2_test
Sandman dream multitask v2 — test split
The test split for fine-tuning
Sandman's on-device dream-analysis model (v2).
See sandman-dream_multitask_v2_train
for the full description of the three tasks (summarize, extract symbols,
interpret a symbol) and the source data.
sandman-dream_multitask_v2_train
Sandman dream multitask v2 — train split
17,300 instruction-following examples for fine-tuning Sandman's on-device
dream-analysis model, built from
sandman-dreambank-v2.
Every row is a single-turn conversation (messages) covering one of three
tasks:
Summarize — read a dream, return a one- or two-sentence summary as JSON.
Extract symbols — return only the concrete nouns literally present in
the dream text, as a JSON array, with an explicit instruction not to
infer or add… See the full description on the dataset page: https://huggingface.co/datasets/mujo-labs/sandman-dream_multitask_v2_train.sandman-dream_multitask_v2_val
Sandman dream multitask v2 — val split
The val split for fine-tuning
Sandman's on-device dream-analysis model (v2).
See sandman-dream_multitask_v2_train
for the full description of the three tasks (summarize, extract symbols,
interpret a symbol) and the source data.
ctms-multitask-sft-v3
CTMS Multi-Task SFT — V3 (uppercase-Snowflake)
Supervised fine-tuning corpus for a Clinical Trial Management System (CTMS) analytics assistant, spanning
7 tasks over a 122-table CTMS schema. This is the V3 build: all SQL uses unquoted identifiers
that resolve against the uppercase-identifier Snowflake schema DUMMY_FORTREA_AI_MODEL.FORTREA_AI_MODEL_V3_CAP.
Data is fully synthetic (generated from a CTMS data generator). It contains no real patient,
investigator, or trial data.… See the full description on the dataset page: https://huggingface.co/datasets/persistent-fm/ctms-multitask-sft-v3.
