datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
self_instructThis dataset splits the original Self-instruct dataset into training (90%) and test (10%).
Qwen2.5-7B-Instruct-Self-Calibration
Efficient Test-Time Scaling via Self-Calibration
This repository contains datasets used in the paper Efficient Test-Time Scaling via Self-Calibration. The datasets are used to evaluate the effectiveness of test-time scaling methods for LLMs. Each config_name in the metadata refers to a different reasoning dataset. More detailed descriptions of each dataset are needed. Consider adding a section for each config_name with a description, statistics, and any other relevant… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Qwen2.5-7B-Instruct-Self-Calibration.GPT-4-Self-Instruct-GermanHere we share a German dataset synthesized using the OpenAI GPT-4 model with Self-Instruct, utilizing some excess Azure credits. Please feel free to use it. All questions and answers are newly generated by GPT-4, without specialized verification, only simple filtering and strict semantic similarity control have been applied.
We hope that this will be helpful for fine-tuning open-source models for non-English languages, particularly German. This dataset will be updated continuously.
hinglish_self_instruct_v0
Hinglish Instruct Dataset using Self Instruct method
The prompt used for generating the samples:
You are asked to come up with a set of 50 diverse task instructions in Hinglish or Hindi.
These task instructions will be given to a GPT model and we will evaluate the GPT model for completing the instructions.
Here are the requirements:
1. Try not to repeat the verb for each instruction to maximize diversity.
2. The language used for the instruction also should be diverse. For example… See the full description on the dataset page: https://huggingface.co/datasets/smangrul/hinglish_self_instruct_v0.self-rewarding_instruct
self-rewarding_instruct
self-rewarding_instructは、
Stanford Alpacaの手法
kunishou/oasst1-89k-jaをseed tasksとして
mistralai/Mixtral-8x22B-Instruct-v0.1
で作った合成データ(Synthetic data)です。モデルの利用にはDeepinfraを利用しています。データセットはmistralai/Mixtral-8x22B-Instruct-v0.1によって精査されています。
self-rewarding用に作成したため、output_exampleとなっていますがInstruction Tuningにも用いれると思います。
Dataset Details
Dataset Description
Curated by: HachiMLLanguage(s) (NLP): Japanese
License: Apache 2.0
Github:… See the full description on the dataset page: https://huggingface.co/datasets/HachiML/self-rewarding_instruct.Self-Instruct-Qwen2.5-72B-Instruct-60k
Self-Instruct-Qwen2.5-72B-Instruct-60k
概要
以下の手順で作成した約6万件の日本語の合成instructionデータセットです。
MagpieとEvol-Instructを使って作成された合成instructionデータであるAratako/Magpie-Tanuki-Instruction-Selected-Evolved-26.5kに対し、Qwen/Qwen2.5-72B-Instruct-GPTQ-Int8を使ってinstructionのカテゴリを付与
同じカテゴリに分類された3つのinstructionをseed taskとしてQwen/Qwen2.5-72B-Instruct-GPTQ-Int8に与え、Self-Instructの手法で7個のinstructionを生成
元論文の実装とは一部異なります。
generated_instruction列が生成されたinstructionです。
ライセンス
基本的にはApache 2.0に準じますが、Qwen… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Self-Instruct-Qwen2.5-72B-Instruct-60k.bigcode_self-oss-instruct-sc2-exec-filter-50k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bigcode_self-oss-instruct-sc2-exec-filter-50k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bigcode/self-oss-instruct-sc2-exec-filter-50k with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/bigcode_self-oss-instruct-sc2-exec-filter-50k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.Self-Instruct-Japanese-Elzya-13BA Japanese dataset generated with an opensource elyza/ELYZA-japanese-Llama-2-13b-instruct model.
This dataset is used in evaluating AI-generated text detection methods and is well-suited for self-instruct methods.
The instructions were taken from:
https://huggingface.co/datasets/CausalLM/GPT-4-Self-Instruct-Japanese
The model used is:
https://huggingface.co/elyza/ELYZA-japanese-Llama-2-7b
License: refer to the model's license
Self-Instruct-Japanese-Qwen1.5-14BA Japanese dataset generated with Qwen/Qwen1.5-14B model.
This dataset is used in evaluating AI-generated text detection methods and is well-suited for self-instruct methods.
The instructions were taken from:
https://huggingface.co/datasets/CausalLM/GPT-4-Self-Instruct-Japanese
The model used is:
https://huggingface.co/Qwen/Qwen1.5-14B
License: Please refer to the model's license
