datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Telugu-MultiTask-Instruct-77K
Telugu MultiTask Instruct 77K — Adaption AutoScientist Challenge Dataset
Powered by Adaptive Data — Adaption Labs
Dataset Description
A large-scale, multi-task Telugu instruction-tuning dataset combining 77,653 rows from 7 open-source Telugu NLP collections. Covers diverse tasks including news summarization, QA, creative writing, translation, and general instruction following — all processed through the Adaption Labs AutoScientist platform for quality… See the full description on the dataset page: https://huggingface.co/datasets/narendarcodes/Telugu-MultiTask-Instruct-77K.sec-extraction-multitask-v4
SEC Extraction Multitask v4
Instruction-tuning dataset for fine-tuning a small language model (e.g. Gemma 4 E2B) to extract structured data from SEC filings across three verticals:
Exhibit 10 (contracts) — financial terms from executive employment, credit agreements, indemnification, licensing, and similar filings
DEF 14A (proxy statements) — executive compensation, governance items, say-on-pay
MD&A (10-K / 10-Q Management's Discussion & Analysis) — operating metrics, segment… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-extraction-multitask-v4.thai-multitask-starter
Thai Multitask 9.6K
ชุดข้อมูลตั้งต้นสำหรับ instruction tuning ภาษาไทย ครอบคลุมงานสนทนา ถาม–ตอบ สรุป
แปล จำแนกข้อความ ตรวจแก้ภาษา คณิตศาสตร์ และ structured output
ข้อมูลทุกแถวสร้างขึ้นใหม่ด้วยกฎแบบ deterministic ไม่มีการคัดลอกจากเว็บไซต์หรือ
ข้อมูลส่วนบุคคลจริง เหมาะสำหรับทดลอง supervised fine-tuning และทดสอบ pipeline
แต่ควรเพิ่มข้อมูลที่มนุษย์ตรวจทานและข้อมูลภาษาธรรมชาติก่อนใช้กับระบบจริง
จำนวนข้อมูลทั้งหมด 9,599 ตัวอย่าง: train 8,639, validation 480 และ test 480… See the full description on the dataset page: https://huggingface.co/datasets/Phettae/thai-multitask-starter.sandman-dream_multitask_v2_train
Sandman dream multitask v2 — train split
17,300 instruction-following examples for fine-tuning Sandman's on-device
dream-analysis model, built from
sandman-dreambank-v2.
Every row is a single-turn conversation (messages) covering one of three
tasks:
Summarize — read a dream, return a one- or two-sentence summary as JSON.
Extract symbols — return only the concrete nouns literally present in
the dream text, as a JSON array, with an explicit instruction not to
infer or add… See the full description on the dataset page: https://huggingface.co/datasets/mujo-labs/sandman-dream_multitask_v2_train.multi-task-instructionsandman-dream_multitask_v2_test
Sandman dream multitask v2 — test split
The test split for fine-tuning
Sandman's on-device dream-analysis model (v2).
See sandman-dream_multitask_v2_train
for the full description of the three tasks (summarize, extract symbols,
interpret a symbol) and the source data.
sandman-dream_multitask_v2_val
Sandman dream multitask v2 — val split
The val split for fine-tuning
Sandman's on-device dream-analysis model (v2).
See sandman-dream_multitask_v2_train
for the full description of the three tasks (summarize, extract symbols,
interpret a symbol) and the source data.
ctms-multitask-sft-v3
CTMS Multi-Task SFT — V3 (uppercase-Snowflake)
Supervised fine-tuning corpus for a Clinical Trial Management System (CTMS) analytics assistant, spanning
7 tasks over a 122-table CTMS schema. This is the V3 build: all SQL uses unquoted identifiers
that resolve against the uppercase-identifier Snowflake schema DUMMY_FORTREA_AI_MODEL.FORTREA_AI_MODEL_V3_CAP.
Data is fully synthetic (generated from a CTMS data generator). It contains no real patient,
investigator, or trial data.… See the full description on the dataset page: https://huggingface.co/datasets/persistent-fm/ctms-multitask-sft-v3.
