self-instruct
RyanYr_-_self-correct_Llama-3.2-3B-Instruct_MATH_bon_iter2-ggufRyanYr_-_self-correct_Llama-3.2-3B-Instruct_MATH_0.75-1_bon_iter4-ggufRyanYr_-_self-correct_Llama-3.2-3B-Instruct_MATH_0.5-0.75_bon_iter3-ggufRyanYr_-_self-correct_Llama-3.2-3B-Instruct_MATH_0.25-0.5_bon_iter2-ggufQwQ-32B-Preview-Self-instruct-3x-TIES-v1.0-i1-GGUFRyanYr_-_self-correct_Llama-3.2-3B-Instruct_MATH_0-0.25_bon_iter1-ggufOcelot-Ko-self-instruction-10.8B-v1.0-i1-GGUFRyanYr_-_self-correct_Llama-3.2-3B-Instruct_MATH_bon_iter3-gguf
self-oss-instruct-sc2-exec-filter-50kFinal self-alignment training dataset for StarCoder2-Instruct.
seed: Contains the seed Python function
concepts: Contains the concepts generated from the seed
instruction: Contains the instruction generated from the concepts
response: Contains the execution-validated response to the instruction
This dataset utilizes seed Python functions derived from the MultiPL-T pipeline.
testing_self_instruct_small
Dataset Card for "testing_self_instruct_small"
More Information needed
MSC-Self-Instruct
MemGPT
This is the self-instruct dataset of MSC conversations used for MemGPT paper. For more information please refer to memgpt.ai
The MSC dataset is a multi-round human conversations. In this dataset, our goal is to come up with a conversation opener, that is personalized to the user by referencing topics from the previous conversations.
These were generated while evaluating MemGPT.
self-instruct-starcoder
Self-instruct-starcoder
Summary
Self-instruct-starcoder is a dataset that was generated by prompting starcoder to generate new instructions based on some human-written seed instructions.
The underlying process is explained in the paper self-instruct. This algorithm gave birth to famous machine generated
datasets such as Alpaca and Code Alpaca which are two datasets
obtained by prompting OpenAI text-davinci-003 engine.
Our approach
While our method is… See the full description on the dataset page: https://huggingface.co/datasets/codeparrot/self-instruct-starcoder.self_instructSelf-Instruct is a dataset that contains 52k instructions, paired with 82K instance inputs and outputs. This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction better.Multi-modal-Self-instruct
Dataset Description
Paper Information
Dataset Examples
Leaderboard
Dataset Usage
Data Downloading
Data Format
Evaluation
Citation
You can download the zip dataset directly, and both train and test subsets are collected in Multi-modal-Self-instruct.zip.
Dataset Description
Multi-Modal Self-Instruct dataset utilizes large language models and their code capabilities to synthesize massive abstract images and visual reasoning instructions across daily scenarios. This benchmark… See the full description on the dataset page: https://huggingface.co/datasets/zwq2018/Multi-modal-Self-instruct.
