CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01xingqiang /GPRadar-Defect-MultiTask GPRadar-Defect-MultiTask 数据集 本仓库包含用于微调PaLI-GEMMA多模态模型的地质雷达(GPR)缺陷检测数据集。该数据集专注于地下结构中的空洞和裂缝检测与分析。 数据集结构 数据集组织如下: dataset/ ├── annotations/ - 包含JSON和JSONL格式的标注文件 │ ├── _annotations.train.jsonl - 训练集标注 │ ├── _annotations.valid.jsonl - 验证集标注 │ ├── _annotations.test.jsonl - 测试集标注 │ ├── p-1.v1i.paligemma/ - 主数据集元数据 │ └── p-1.v1i.paligemma-multimodal/ - 多模态数据集元数据 ├── images/ - 包含所有图像文件 特点 包含874张带注释的地质雷达扫描图像 图像预处理为640x640像素大小 支持多种任务类型:缺陷检测、位置定位和描述生成… See the full description on the dataset page: https://huggingface.co/datasets/xingqiang/GPRadar-Defect-MultiTask.imageobject-detection1K<n<10K0 likes148 downloads2y agoHugging Face02LiZHENGzai /GPRadar-Defect-MultiTask GPRadar-Defect-MultiTask 数据集 本仓库包含用于微调PaLI-GEMMA多模态模型的地质雷达(GPR)缺陷检测数据集。该数据集专注于地下结构中的空洞和裂缝检测与分析。 数据集结构 数据集组织如下: dataset/ ├── annotations/ - 包含JSON和JSONL格式的标注文件 │ ├── _annotations.train.jsonl - 训练集标注 │ ├── _annotations.valid.jsonl - 验证集标注 │ ├── _annotations.test.jsonl - 测试集标注 │ ├── p-1.v1i.paligemma/ - 主数据集元数据 │ └── p-1.v1i.paligemma-multimodal/ - 多模态数据集元数据 ├── images/ - 包含所有图像文件 特点 包含874张带注释的地质雷达扫描图像 图像预处理为640x640像素大小 支持多种任务类型:缺陷检测、位置定位和描述生成… See the full description on the dataset page: https://huggingface.co/datasets/LiZHENGzai/GPRadar-Defect-MultiTask.imageobject-detection1K<n<10K0 likes93 downloads6mo agoHugging Face03TaskPuppyAI /lunamax-multitask-programming-1000 LunaMax Multitask Programming 1000 A 1,000-record synthetic multitask programming dataset generated with ChatGPT LunaMax. The recovered dataset combines code review, implementation, bug and severity classification, and strict output-contract tasks across multiple programming languages. The historical source shards were reviewed with ChatGPT 5.6 Sol High according to dataset creator confirmation. During Hugging Face publication preparation, all 1,000 records received a new… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/lunamax-multitask-programming-1000.text1K<n<10K0 likes69 downloads16d agoHugging Face04TaskPuppyAI /lunamax-multitask-programming-250 LunaMax Multitask Programming 250 A 250-record synthetic multitask programming dataset generated with ChatGPT LunaMax. The dataset combines structured and free-form code review, implementation, bug and severity classification, and strict output-contract tasks across multiple programming languages. Generation and historical-review attribution are based on dataset creator confirmation. Dataset Summary The publication dataset contains: 250 records 250 unique records… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/lunamax-multitask-programming-250.textn<1K0 likes47 downloads16d agoHugging Face05hamishivi /rds-sels-multitask-rrmax-top326k RDS+ Selected Multitask 326k This is the dataset (and associated scores) selected by RDS+ when selecting 326k samples for multiple tasks at once. For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning. This was used to train this model. This dataset is selected from Tulu 2 unfiltered, and please see that page for more information on sources. License We are releasing this dataset under the terms of ODC-BY. By using this, you… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-multitask-rrmax-top326k.text100K<n<1M1 likes46 downloads2y agoHugging Face06narendarcodes /Telugu-MultiTask-Instruct-77K Telugu MultiTask Instruct 77K — Adaption AutoScientist Challenge Dataset Powered by Adaptive Data — Adaption Labs Dataset Description A large-scale, multi-task Telugu instruction-tuning dataset combining 77,653 rows from 7 open-source Telugu NLP collections. Covers diverse tasks including news summarization, QA, creative writing, translation, and general instruction following — all processed through the Adaption Labs AutoScientist platform for quality… See the full description on the dataset page: https://huggingface.co/datasets/narendarcodes/Telugu-MultiTask-Instruct-77K.textquestion-answering10K<n<100K1 likes43 downloads3mo agoHugging Face07TheTokenFactory /sec-extraction-multitask-v4 SEC Extraction Multitask v4 Instruction-tuning dataset for fine-tuning a small language model (e.g. Gemma 4 E2B) to extract structured data from SEC filings across three verticals: Exhibit 10 (contracts) — financial terms from executive employment, credit agreements, indemnification, licensing, and similar filings DEF 14A (proxy statements) — executive compensation, governance items, say-on-pay MD&A (10-K / 10-Q Management's Discussion & Analysis) — operating metrics, segment… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-extraction-multitask-v4.texttext-generation1K<n<10K0 likes40 downloads5mo agoHugging Face08Kushalkhemka /cybersec-chatml-multitask-v1 Cybersecurity ChatML Multitask Dataset (v1) Combined split for both detection and patch tasks. Files chatml_multitask_train.jsonl chatml_multitask_val.jsonl chatml_build_manifest.json unsloth_best_params_glm47flash_multitask.json Output format Detection samples: strict JSON schema output Patch samples: patched code only text100K<n<1M0 likes39 downloads6mo agoHugging Face09PJMixers /vicgalle_configurable-system-prompt-multitask-PreferenceShareGPTtextreinforcement-learning1K<n<10K5 likes37 downloads2y agoHugging Face10Phettae /thai-multitask-starter Thai Multitask 9.6K ชุดข้อมูลตั้งต้นสำหรับ instruction tuning ภาษาไทย ครอบคลุมงานสนทนา ถาม–ตอบ สรุป แปล จำแนกข้อความ ตรวจแก้ภาษา คณิตศาสตร์ และ structured output ข้อมูลทุกแถวสร้างขึ้นใหม่ด้วยกฎแบบ deterministic ไม่มีการคัดลอกจากเว็บไซต์หรือ ข้อมูลส่วนบุคคลจริง เหมาะสำหรับทดลอง supervised fine-tuning และทดสอบ pipeline แต่ควรเพิ่มข้อมูลที่มนุษย์ตรวจทานและข้อมูลภาษาธรรมชาติก่อนใช้กับระบบจริง จำนวนข้อมูลทั้งหมด 9,599 ตัวอย่าง: train 8,639, validation 480 และ test 480… See the full description on the dataset page: https://huggingface.co/datasets/Phettae/thai-multitask-starter.texttext-generation1K<n<10K0 likes33 downloads1mo agoHugging Face11AethronPhantom /nexa-science-multitask-balanced Nexa Science Multitask Balanced This dataset is a curated, instruction-formatted scientific multitask mixture for: claim verification (<TASK:VERIFY>) abstract-grounded biomedical QA (<TASK:QA>) retrieval relevance re-ranking (<TASK:RERANK>) Format Each row is JSONL with: {task, instruction, input, output, meta} Splits Included train_balanced_short.jsonl val_balanced_short.jsonl stats_balanced_short.json Notes QA in this balanced release is… See the full description on the dataset page: https://huggingface.co/datasets/AethronPhantom/nexa-science-multitask-balanced.texttext-classification10K<n<100K0 likes31 downloads7mo agoHugging Face12cle-13 /rutooro_multitask Rutooro Multitask Dataset This dataset contains a collection of instruction-response pairs for fine-tuning a Large Language Model (LLM) on the Rutooro language. The dataset is prepared for a multi-task learning approach, including: Translation: English to Rutooro. Monolingual Generation: Continued stories and prose in Rutooro. Grammar Instructions: Explanations of Rutooro grammar rules. Data Source The data was sourced from [mention your source, e.g., "manual… See the full description on the dataset page: https://huggingface.co/datasets/cle-13/rutooro_multitask.text1K<n<10K0 likes23 downloads1y agoHugging Face13nmd2k /multi-task-instructiontexttext-generation100K<n<1M0 likes19 downloads3y agoHugging Face14hamishivi /rds-sels-tulu-3-multitask-rrmax-939k RDS+ Selected Tulu 3 Multitask 939k This is the dataset (and associated scores) selected by RDS+ when selecting 939k samples targeting multiple downstream tasks. For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning. This was used to train this model. This dataset is selected from Tulu 3 unfiltered, and please see that page for more information on sources. License This dataset is licensed under ODC-BY-1.0. It is intended… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-tulu-3-multitask-rrmax-939k.text100K<n<1M0 likes18 downloads2y agoHugging Face15persistent-fm /ctms-multitask-sft-v6gated CTMS Multi-task SFT — V6 A matched pair of corpora for a clinical-trial-management text-to-SQL agent, differing in exactly one variable: whether generate_sql rows carry a <think> reasoning trace. run_a (control) run_b (traced) total 23,049 23,049 train / val / test 18,698 / 2,172 / 2,179 18,698 / 2,172 / 2,179 traced train SQL rows 0 10,125 (81.0%) gold SQL identical, byte-for-byte identical, byte-for-byte Tasks task n… See the full description on the dataset page: https://huggingface.co/datasets/persistent-fm/ctms-multitask-sft-v6.text10K<n<100K0 likes18 downloads1mo agoHugging Face16lohoz /Smart-Contract-MultiTask-Dataset Overview This is a dataset designed for smart contract generation. It includes two subsets: Requirement-FSM-Code subset: Contains user requirement descriptions, finite state machine (FSM) representations, and corresponding smart contract code. Comment-Code subset: Includes functional comments and their corresponding implementation code. Dataset Structure Subset 1: Requirement-FSM-Code Description: Contains natural language descriptions of user requirements… See the full description on the dataset page: https://huggingface.co/datasets/lohoz/Smart-Contract-MultiTask-Dataset.text10K<n<100K0 likes17 downloads2y agoHugging Face17mujo-labs /sandman-dream_multitask_v2_test Sandman dream multitask v2 — test split The test split for fine-tuning Sandman's on-device dream-analysis model (v2). See sandman-dream_multitask_v2_train for the full description of the three tasks (summarize, extract symbols, interpret a symbol) and the source data. texttext-generation1K<n<10K0 likes16 downloads3d agoHugging Face18mujo-labs /sandman-dream_multitask_v2_train Sandman dream multitask v2 — train split 17,300 instruction-following examples for fine-tuning Sandman's on-device dream-analysis model, built from sandman-dreambank-v2. Every row is a single-turn conversation (messages) covering one of three tasks: Summarize — read a dream, return a one- or two-sentence summary as JSON. Extract symbols — return only the concrete nouns literally present in the dream text, as a JSON array, with an explicit instruction not to infer or add… See the full description on the dataset page: https://huggingface.co/datasets/mujo-labs/sandman-dream_multitask_v2_train.texttext-generation10K<n<100K0 likes15 downloads3d agoHugging Face19mujo-labs /sandman-dream_multitask_v2_val Sandman dream multitask v2 — val split The val split for fine-tuning Sandman's on-device dream-analysis model (v2). See sandman-dream_multitask_v2_train for the full description of the three tasks (summarize, extract symbols, interpret a symbol) and the source data. texttext-generation1K<n<10K0 likes14 downloads3d agoHugging Face20CL-From-Nothing /rlve-multitask-qwen3-4b-rollouts-n4-tokens16384tabular1K<n<10K0 likes10 downloads5mo agoHugging Face21samirmsallem /wiki_definitions_de_multitask Dataset Card for Wikipedia Definitions for Multitask (NER/Text Classification) The Wikipedia Definitions for Multitask (NER/Text Classification) dataset is a dataset to train language models to recognize definition sentences and non-definition sentences. The dataset includes training and test data to recognize this discipline by Named Entity Recognition, but also by Sentence Classification. Dataset Sources Wikimedia/wikipedia Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/samirmsallem/wiki_definitions_de_multitask.texttext-classification10K<n<100K0 likes8 downloads1y agoHugging Face22GilbertAkham /gilbert-multitask-mix DATASET_README.md --- language: - en task_categories: - text-generation - summarization - question-answering - conversational tags: - multitask - email - stories - qa - summarization - chat license: - cc-by-4.0 - apache-2.0 - mit --- # Gilbert-Multitask-Mix A diverse multitask dataset for text generation training, combining samples from 5 different domains with structured prompt formatting. ## Dataset Description This dataset contains 6,500+ examples across multiple text… See the full description on the dataset page: https://huggingface.co/datasets/GilbertAkham/gilbert-multitask-mix.text100K<n<1M0 likes6 downloads11mo agoHugging Face23CL-From-Nothing /rlve-multitask-qwen3-4b-n4-randcut512-4096x20-completed-by-qwen3-4b-thinking-r16384tabular10K<n<100K0 likes5 downloads5mo agoHugging Face24persistent-fm /ctms-multitask-sft-v3gated CTMS Multi-Task SFT — V3 (uppercase-Snowflake) Supervised fine-tuning corpus for a Clinical Trial Management System (CTMS) analytics assistant, spanning 7 tasks over a 122-table CTMS schema. This is the V3 build: all SQL uses unquoted identifiers that resolve against the uppercase-identifier Snowflake schema DUMMY_FORTREA_AI_MODEL.FORTREA_AI_MODEL_V3_CAP. Data is fully synthetic (generated from a CTMS data generator). It contains no real patient, investigator, or trial data.… See the full description on the dataset page: https://huggingface.co/datasets/persistent-fm/ctms-multitask-sft-v3.texttext-generation10K<n<100K0 likes4 downloads1mo agoHugging Face25renpley2 /raij-instruct-multitasktextn<1K0 likes3 downloads1y agoHugging Face26persistent-fm /ctms-multitask-sft-v10-2gatedtext100K<n<1M0 likes3 downloads15d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.