datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
self-oss-instruct-sc2-exec-filter-50kFinal self-alignment training dataset for StarCoder2-Instruct.
seed: Contains the seed Python function
concepts: Contains the concepts generated from the seed
instruction: Contains the instruction generated from the concepts
response: Contains the execution-validated response to the instruction
This dataset utilizes seed Python functions derived from the MultiPL-T pipeline.
testing_self_instruct_small
Dataset Card for "testing_self_instruct_small"
More Information needed
MSC-Self-Instruct
MemGPT
This is the self-instruct dataset of MSC conversations used for MemGPT paper. For more information please refer to memgpt.ai
The MSC dataset is a multi-round human conversations. In this dataset, our goal is to come up with a conversation opener, that is personalized to the user by referencing topics from the previous conversations.
These were generated while evaluating MemGPT.
self-instruct-starcoder
Self-instruct-starcoder
Summary
Self-instruct-starcoder is a dataset that was generated by prompting starcoder to generate new instructions based on some human-written seed instructions.
The underlying process is explained in the paper self-instruct. This algorithm gave birth to famous machine generated
datasets such as Alpaca and Code Alpaca which are two datasets
obtained by prompting OpenAI text-davinci-003 engine.
Our approach
While our method is… See the full description on the dataset page: https://huggingface.co/datasets/codeparrot/self-instruct-starcoder.self_instructSelf-Instruct is a dataset that contains 52k instructions, paired with 82K instance inputs and outputs. This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction better.Multi-modal-Self-instruct
Dataset Description
Paper Information
Dataset Examples
Leaderboard
Dataset Usage
Data Downloading
Data Format
Evaluation
Citation
You can download the zip dataset directly, and both train and test subsets are collected in Multi-modal-Self-instruct.zip.
Dataset Description
Multi-Modal Self-Instruct dataset utilizes large language models and their code capabilities to synthesize massive abstract images and visual reasoning instructions across daily scenarios. This benchmark… See the full description on the dataset page: https://huggingface.co/datasets/zwq2018/Multi-modal-Self-instruct.dev_set_v2_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_20260825_134035terminal_bench_2_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_20260f49e33daswebench_verified_random_100_folders_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_c2a08420self-oss-instruct-sc2-H4StarCoder2-Self-Instruct-OSS-50k dataset formatted to be compatible with the alignement-handbook for SFT.
selfinstruct-naive-sandboxes-2-verified-qwen3.5-122b-131k-opencode-tracesrl__24GPU_shaped__selfinstruct-naive-sandboxes-2-verified__exp_tas_optimal_comb__40-0a3-rl-DCAgent_selfinstruct-naive-sandboxes-2-verifiedbfcl_parity_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_20260604_192648swebench_verified_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_2d90a6650self_instructThis dataset splits the original Self-instruct dataset into training (90%) and test (10%).
medagentbench_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_2026061f18c01terminal_bench_2_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_204c8a44f2self-oss-instruct-sc2-exec-filter-prompt-codes-test-50kswebench_verified_random_100_folders_a3_rl_DCAgent_selfinstruct_naive_sandboxes6b437cbaLlama_3.1-8B-Instruct-Self-CalibrationThe official repository which contains the code and pre-trained models/datasets for our paper Efficient Test-Time Scaling via Self-Calibration.
🔥 Updates
[2025-3-3]: We released our paper.
[2025-2-25]: We released our codes, models and datasets.
🏴 Overview
We propose an efficient test-time scaling method by using model confidence for dynamically sampling adjustment, since confidence can be seen as an intrinsic measure that directly reflects model… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Llama_3.1-8B-Instruct-Self-Calibration.gaia_127_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_20260604_192659dev_set_v2_rl__24GPU_shaped__selfinstruct_naive_sandboxes_2_verified__exp_tas_o57316c9adev_set_v2_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_20260604_192604multimodal-vqa-self-instruct-enriched
Multimodal VQA – Self-Instruct-enriched
Overview
This dataset is an enriched, cleaned, and metadata-enhanced version of zwq2018/Multi-modal-Self-instruct.It pairs images with natural language questions and answers, making it ideal for Vision-Language Model (VLM) training, benchmarking, and instruction tuning.
Dataset Summary
Total samples: 75,000+ (64,796 train, 11,193 test)
Modalities: Image + Text (Questions) + Text (Answers)
Task Types: Visual Question… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/multimodal-vqa-self-instruct-enriched.x-self-instruct-seed-32
Dataset Card for xOA22 - Multilingual Prompts from OpenAssistant
Dataset Summary
x-self-instruct-seed-32 consists of 32 prompts chosen out of the 252 prompts in the self-instruct-seed dataset from the Self-Instruct paper. These 32 prompts were filtered out according to the following criteria:
Should be natural in a chat setting
Therefore, we filter out any prompts with "few-shot examples", as these are all instruction prompts that we consider unnatural in a chat setting… See the full description on the dataset page: https://huggingface.co/datasets/sambanovasystems/x-self-instruct-seed-32.aider_polyglot_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_202690541bceterminal_bench_2_rl__24GPU_shaped__selfinstruct_naive_sandboxes_2_verified__exp93c24543financeagent_terminal_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_5349063bself-instruct-safety-alignment[EMNLP 2024] Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
🌐 Homepage | 📖 Paper | 🤗 Dataset (Data Advisor) | 🤗 Dataset (Self-Instruct)
Disclaimer
The dataset contains content that may be offensive or harmful. This dataset is intended for research purposes, specifically to support efforts aimed at creating safer and less harmful AI systems. Please engage with it responsibly and at your own risk.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/fwnlp/self-instruct-safety-alignment.
