datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPRadar-Defect-MultiTask
GPRadar-Defect-MultiTask 数据集
本仓库包含用于微调PaLI-GEMMA多模态模型的地质雷达(GPR)缺陷检测数据集。该数据集专注于地下结构中的空洞和裂缝检测与分析。
数据集结构
数据集组织如下:
dataset/
├── annotations/ - 包含JSON和JSONL格式的标注文件
│ ├── _annotations.train.jsonl - 训练集标注
│ ├── _annotations.valid.jsonl - 验证集标注
│ ├── _annotations.test.jsonl - 测试集标注
│ ├── p-1.v1i.paligemma/ - 主数据集元数据
│ └── p-1.v1i.paligemma-multimodal/ - 多模态数据集元数据
├── images/ - 包含所有图像文件
特点
包含874张带注释的地质雷达扫描图像
图像预处理为640x640像素大小
支持多种任务类型:缺陷检测、位置定位和描述生成… See the full description on the dataset page: https://huggingface.co/datasets/xingqiang/GPRadar-Defect-MultiTask.GPRadar-Defect-MultiTask
GPRadar-Defect-MultiTask 数据集
本仓库包含用于微调PaLI-GEMMA多模态模型的地质雷达(GPR)缺陷检测数据集。该数据集专注于地下结构中的空洞和裂缝检测与分析。
数据集结构
数据集组织如下:
dataset/
├── annotations/ - 包含JSON和JSONL格式的标注文件
│ ├── _annotations.train.jsonl - 训练集标注
│ ├── _annotations.valid.jsonl - 验证集标注
│ ├── _annotations.test.jsonl - 测试集标注
│ ├── p-1.v1i.paligemma/ - 主数据集元数据
│ └── p-1.v1i.paligemma-multimodal/ - 多模态数据集元数据
├── images/ - 包含所有图像文件
特点
包含874张带注释的地质雷达扫描图像
图像预处理为640x640像素大小
支持多种任务类型:缺陷检测、位置定位和描述生成… See the full description on the dataset page: https://huggingface.co/datasets/LiZHENGzai/GPRadar-Defect-MultiTask.lunamax-multitask-programming-1000
LunaMax Multitask Programming 1000
A 1,000-record synthetic multitask programming dataset generated with
ChatGPT LunaMax.
The recovered dataset combines code review, implementation, bug and severity
classification, and strict output-contract tasks across multiple programming
languages.
The historical source shards were reviewed with ChatGPT 5.6 Sol High according
to dataset creator confirmation. During Hugging Face publication preparation,
all 1,000 records received a new… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/lunamax-multitask-programming-1000.lunamax-multitask-programming-250
LunaMax Multitask Programming 250
A 250-record synthetic multitask programming dataset generated with
ChatGPT LunaMax.
The dataset combines structured and free-form code review, implementation,
bug and severity classification, and strict output-contract tasks across
multiple programming languages.
Generation and historical-review attribution are based on dataset creator
confirmation.
Dataset Summary
The publication dataset contains:
250 records
250 unique records… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/lunamax-multitask-programming-250.rds-sels-multitask-rrmax-top326k
RDS+ Selected Multitask 326k
This is the dataset (and associated scores) selected by RDS+ when selecting 326k samples for multiple tasks at once.
For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning.
This was used to train this model.
This dataset is selected from Tulu 2 unfiltered, and please see that page for more information on sources.
License
We are releasing this dataset under the terms of ODC-BY. By using this, you… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-multitask-rrmax-top326k.Telugu-MultiTask-Instruct-77K
Telugu MultiTask Instruct 77K — Adaption AutoScientist Challenge Dataset
Powered by Adaptive Data — Adaption Labs
Dataset Description
A large-scale, multi-task Telugu instruction-tuning dataset combining 77,653 rows from 7 open-source Telugu NLP collections. Covers diverse tasks including news summarization, QA, creative writing, translation, and general instruction following — all processed through the Adaption Labs AutoScientist platform for quality… See the full description on the dataset page: https://huggingface.co/datasets/narendarcodes/Telugu-MultiTask-Instruct-77K.sec-extraction-multitask-v4
SEC Extraction Multitask v4
Instruction-tuning dataset for fine-tuning a small language model (e.g. Gemma 4 E2B) to extract structured data from SEC filings across three verticals:
Exhibit 10 (contracts) — financial terms from executive employment, credit agreements, indemnification, licensing, and similar filings
DEF 14A (proxy statements) — executive compensation, governance items, say-on-pay
MD&A (10-K / 10-Q Management's Discussion & Analysis) — operating metrics, segment… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-extraction-multitask-v4.cybersec-chatml-multitask-v1
Cybersecurity ChatML Multitask Dataset (v1)
Combined split for both detection and patch tasks.
Files
chatml_multitask_train.jsonl
chatml_multitask_val.jsonl
chatml_build_manifest.json
unsloth_best_params_glm47flash_multitask.json
Output format
Detection samples: strict JSON schema output
Patch samples: patched code only
vicgalle_configurable-system-prompt-multitask-PreferenceShareGPTthai-multitask-starter
Thai Multitask 9.6K
ชุดข้อมูลตั้งต้นสำหรับ instruction tuning ภาษาไทย ครอบคลุมงานสนทนา ถาม–ตอบ สรุป
แปล จำแนกข้อความ ตรวจแก้ภาษา คณิตศาสตร์ และ structured output
ข้อมูลทุกแถวสร้างขึ้นใหม่ด้วยกฎแบบ deterministic ไม่มีการคัดลอกจากเว็บไซต์หรือ
ข้อมูลส่วนบุคคลจริง เหมาะสำหรับทดลอง supervised fine-tuning และทดสอบ pipeline
แต่ควรเพิ่มข้อมูลที่มนุษย์ตรวจทานและข้อมูลภาษาธรรมชาติก่อนใช้กับระบบจริง
จำนวนข้อมูลทั้งหมด 9,599 ตัวอย่าง: train 8,639, validation 480 และ test 480… See the full description on the dataset page: https://huggingface.co/datasets/Phettae/thai-multitask-starter.nexa-science-multitask-balanced
Nexa Science Multitask Balanced
This dataset is a curated, instruction-formatted scientific multitask mixture for:
claim verification (<TASK:VERIFY>)
abstract-grounded biomedical QA (<TASK:QA>)
retrieval relevance re-ranking (<TASK:RERANK>)
Format
Each row is JSONL with:
{task, instruction, input, output, meta}
Splits Included
train_balanced_short.jsonl
val_balanced_short.jsonl
stats_balanced_short.json
Notes
QA in this balanced release is… See the full description on the dataset page: https://huggingface.co/datasets/AethronPhantom/nexa-science-multitask-balanced.rutooro_multitask
Rutooro Multitask Dataset
This dataset contains a collection of instruction-response pairs for fine-tuning a Large Language Model (LLM) on the Rutooro language. The dataset is prepared for a multi-task learning approach, including:
Translation: English to Rutooro.
Monolingual Generation: Continued stories and prose in Rutooro.
Grammar Instructions: Explanations of Rutooro grammar rules.
Data Source
The data was sourced from [mention your source, e.g., "manual… See the full description on the dataset page: https://huggingface.co/datasets/cle-13/rutooro_multitask.multi-task-instructionrds-sels-tulu-3-multitask-rrmax-939k
RDS+ Selected Tulu 3 Multitask 939k
This is the dataset (and associated scores) selected by RDS+ when selecting 939k samples targeting multiple downstream tasks.
For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning.
This was used to train this model.
This dataset is selected from Tulu 3 unfiltered, and please see that page for more information on sources.
License
This dataset is licensed under ODC-BY-1.0. It is intended… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-tulu-3-multitask-rrmax-939k.ctms-multitask-sft-v6
CTMS Multi-task SFT — V6
A matched pair of corpora for a clinical-trial-management text-to-SQL agent, differing in
exactly one variable: whether generate_sql rows carry a <think> reasoning trace.
run_a (control)
run_b (traced)
total
23,049
23,049
train / val / test
18,698 / 2,172 / 2,179
18,698 / 2,172 / 2,179
traced train SQL rows
0
10,125 (81.0%)
gold SQL
identical, byte-for-byte
identical, byte-for-byte
Tasks
task
n… See the full description on the dataset page: https://huggingface.co/datasets/persistent-fm/ctms-multitask-sft-v6.Smart-Contract-MultiTask-Dataset
Overview
This is a dataset designed for smart contract generation. It includes two subsets:
Requirement-FSM-Code subset: Contains user requirement descriptions, finite state machine (FSM) representations, and corresponding smart contract code.
Comment-Code subset: Includes functional comments and their corresponding implementation code.
Dataset Structure
Subset 1: Requirement-FSM-Code
Description: Contains natural language descriptions of user requirements… See the full description on the dataset page: https://huggingface.co/datasets/lohoz/Smart-Contract-MultiTask-Dataset.humanoid-multi-step-task-instructions
Humanoid Multi-Step Task Instructions
A structured dataset containing multi-step task instructions for humanoid robots.
Use Cases
Task planning
Autonomous execution
Robotics simulation
sandman-dream_multitask_v2_test
Sandman dream multitask v2 — test split
The test split for fine-tuning
Sandman's on-device dream-analysis model (v2).
See sandman-dream_multitask_v2_train
for the full description of the three tasks (summarize, extract symbols,
interpret a symbol) and the source data.
sandman-dream_multitask_v2_train
Sandman dream multitask v2 — train split
17,300 instruction-following examples for fine-tuning Sandman's on-device
dream-analysis model, built from
sandman-dreambank-v2.
Every row is a single-turn conversation (messages) covering one of three
tasks:
Summarize — read a dream, return a one- or two-sentence summary as JSON.
Extract symbols — return only the concrete nouns literally present in
the dream text, as a JSON array, with an explicit instruction not to
infer or add… See the full description on the dataset page: https://huggingface.co/datasets/mujo-labs/sandman-dream_multitask_v2_train.sandman-dream_multitask_v2_val
Sandman dream multitask v2 — val split
The val split for fine-tuning
Sandman's on-device dream-analysis model (v2).
See sandman-dream_multitask_v2_train
for the full description of the three tasks (summarize, extract symbols,
interpret a symbol) and the source data.
rlve-multitask-qwen3-4b-rollouts-n4-tokens16384wiki_definitions_de_multitask
Dataset Card for Wikipedia Definitions for Multitask (NER/Text Classification)
The Wikipedia Definitions for Multitask (NER/Text Classification) dataset is a dataset to train language models to recognize definition sentences and non-definition sentences.
The dataset includes training and test data to recognize this discipline by Named Entity Recognition, but also by Sentence Classification.
Dataset Sources
Wikimedia/wikipedia Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/samirmsallem/wiki_definitions_de_multitask.gilbert-multitask-mix
DATASET_README.md
---
language:
- en
task_categories:
- text-generation
- summarization
- question-answering
- conversational
tags:
- multitask
- email
- stories
- qa
- summarization
- chat
license:
- cc-by-4.0
- apache-2.0
- mit
---
# Gilbert-Multitask-Mix
A diverse multitask dataset for text generation training, combining samples from 5 different domains with structured prompt formatting.
## Dataset Description
This dataset contains 6,500+ examples across multiple text… See the full description on the dataset page: https://huggingface.co/datasets/GilbertAkham/gilbert-multitask-mix.rlve-multitask-qwen3-4b-n4-randcut512-4096x20-completed-by-qwen3-4b-thinking-r16384ctms-multitask-sft-v3
CTMS Multi-Task SFT — V3 (uppercase-Snowflake)
Supervised fine-tuning corpus for a Clinical Trial Management System (CTMS) analytics assistant, spanning
7 tasks over a 122-table CTMS schema. This is the V3 build: all SQL uses unquoted identifiers
that resolve against the uppercase-identifier Snowflake schema DUMMY_FORTREA_AI_MODEL.FORTREA_AI_MODEL_V3_CAP.
Data is fully synthetic (generated from a CTMS data generator). It contains no real patient,
investigator, or trial data.… See the full description on the dataset page: https://huggingface.co/datasets/persistent-fm/ctms-multitask-sft-v3.raij-instruct-multitaskctms-multitask-sft-v10-2
