CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01narendarcodes /Telugu-MultiTask-Instruct-77K Telugu MultiTask Instruct 77K — Adaption AutoScientist Challenge Dataset Powered by Adaptive Data — Adaption Labs Dataset Description A large-scale, multi-task Telugu instruction-tuning dataset combining 77,653 rows from 7 open-source Telugu NLP collections. Covers diverse tasks including news summarization, QA, creative writing, translation, and general instruction following — all processed through the Adaption Labs AutoScientist platform for quality… See the full description on the dataset page: https://huggingface.co/datasets/narendarcodes/Telugu-MultiTask-Instruct-77K.textquestion-answering10K<n<100K1 likes49 downloads3mo agoHugging Face02TheTokenFactory /sec-extraction-multitask-v4 SEC Extraction Multitask v4 Instruction-tuning dataset for fine-tuning a small language model (e.g. Gemma 4 E2B) to extract structured data from SEC filings across three verticals: Exhibit 10 (contracts) — financial terms from executive employment, credit agreements, indemnification, licensing, and similar filings DEF 14A (proxy statements) — executive compensation, governance items, say-on-pay MD&A (10-K / 10-Q Management's Discussion & Analysis) — operating metrics, segment… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-extraction-multitask-v4.texttext-generation1K<n<10K0 likes37 downloads5mo agoHugging Face03Phettae /thai-multitask-starter Thai Multitask 9.6K ชุดข้อมูลตั้งต้นสำหรับ instruction tuning ภาษาไทย ครอบคลุมงานสนทนา ถาม–ตอบ สรุป แปล จำแนกข้อความ ตรวจแก้ภาษา คณิตศาสตร์ และ structured output ข้อมูลทุกแถวสร้างขึ้นใหม่ด้วยกฎแบบ deterministic ไม่มีการคัดลอกจากเว็บไซต์หรือ ข้อมูลส่วนบุคคลจริง เหมาะสำหรับทดลอง supervised fine-tuning และทดสอบ pipeline แต่ควรเพิ่มข้อมูลที่มนุษย์ตรวจทานและข้อมูลภาษาธรรมชาติก่อนใช้กับระบบจริง จำนวนข้อมูลทั้งหมด 9,599 ตัวอย่าง: train 8,639, validation 480 และ test 480… See the full description on the dataset page: https://huggingface.co/datasets/Phettae/thai-multitask-starter.texttext-generation1K<n<10K0 likes31 downloads1mo agoHugging Face04AethronPhantom /nexa-science-multitask-balanced Nexa Science Multitask Balanced This dataset is a curated, instruction-formatted scientific multitask mixture for: claim verification (<TASK:VERIFY>) abstract-grounded biomedical QA (<TASK:QA>) retrieval relevance re-ranking (<TASK:RERANK>) Format Each row is JSONL with: {task, instruction, input, output, meta} Splits Included train_balanced_short.jsonl val_balanced_short.jsonl stats_balanced_short.json Notes QA in this balanced release is… See the full description on the dataset page: https://huggingface.co/datasets/AethronPhantom/nexa-science-multitask-balanced.texttext-classification10K<n<100K0 likes25 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.