Instruction Tuning
Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.VLA_Instruction_TuningThis repository contains the VLA-IT dataset, a curated 650K-sample Vision-Language-Action Instruction Tuning dataset, and the SimplerEnv-Instruct benchmark. These are presented in the paper InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation. The dataset is designed to enable robots to integrate multimodal reasoning with precise action generation, preserving the flexible reasoning of large vision-language models while delivering leading manipulation… See the full description on the dataset page: https://huggingface.co/datasets/ShuaiYang03/VLA_Instruction_Tuning.DST_Multiwoz21_instruction_Tuning
Dataset Card for "DST_Multiwoz21_instruction_tuning"
More Information needed
Instruction_TuningFiles Contents Details :
Post-Process Code Info :
data_process.py
iamai_seed_tasks_v1.csv :
IAMAI's seed tasks - Version 1 (879)
Total Dataset Size : 879
===============================================================================================
iamai_v1.csv :
Instruction Tuning Dataset collected using seeds from iamai_seed_tasks_v1.csv and ChatGPT API for both prompts and outputs (~248k)
Total Dataset Size : ~248k
iamai_summarization_v1.csv :
Article Summarization dataset (both… See the full description on the dataset page: https://huggingface.co/datasets/iamplus/Instruction_Tuning.DORI-instruction-tuning-dataset
DORI Spatial Reasoning Instruction Dataset
Dataset Description
This dataset contains instruction tuning data for spatial reasoning tasks across multiple question types and visual datasets.
Dataset Structure
Dataset Splits
train: 26,626 samples
test: 6,672 samples
Total: 33,298 samples
Question Types
q1
q2
q3
q4
q5
q6
q7
Source Datasets
3d_future
cityscapes
coco
coco_space_sea
get_3d
jta
kitti
nocs_real
objectron… See the full description on the dataset page: https://huggingface.co/datasets/appledora/DORI-instruction-tuning-dataset.dynamics-of-instruction-tuning
💻 [Github Repo] • 📃 [Paper] • 👀 [Preview]
Update
12/01/23: Corrected ambiguous choices in the validation and test sets of the role-play chat data.
Overview
We introduce DoIT, a collection of over 40k human-curated instruction-output pairs in Chinese. This dataset is organized into ten representative ability categories: (1) STEM subject - Biology, (2) Humanity subject - History, (3) Code Generation, (4) Creative Writing, (5) Language proficiency - Chinese, (6)… See the full description on the dataset page: https://huggingface.co/datasets/ChiyuSONG/dynamics-of-instruction-tuning.
