CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceH4 /self_instructThis dataset splits the original Self-instruct dataset into training (90%) and test (10%). texttext-generation10K<n<100K10 likes158 downloads3y agoHugging Face02HINT-lab /Qwen2.5-7B-Instruct-Self-Calibration Efficient Test-Time Scaling via Self-Calibration This repository contains datasets used in the paper Efficient Test-Time Scaling via Self-Calibration. The datasets are used to evaluate the effectiveness of test-time scaling methods for LLMs. Each config_name in the metadata refers to a different reasoning dataset. More detailed descriptions of each dataset are needed. Consider adding a section for each config_name with a description, statistics, and any other relevant… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Qwen2.5-7B-Instruct-Self-Calibration.tabulartext-generation100K<n<1M0 likes111 downloads2y agoHugging Face03CausalLM /GPT-4-Self-Instruct-GermanHere we share a German dataset synthesized using the OpenAI GPT-4 model with Self-Instruct, utilizing some excess Azure credits. Please feel free to use it. All questions and answers are newly generated by GPT-4, without specialized verification, only simple filtering and strict semantic similarity control have been applied. We hope that this will be helpful for fine-tuning open-source models for non-English languages, particularly German. This dataset will be updated continuously. texttext-generation10K<n<100K30 likes88 downloads2y agoHugging Face04smangrul /hinglish_self_instruct_v0 Hinglish Instruct Dataset using Self Instruct method The prompt used for generating the samples: You are asked to come up with a set of 50 diverse task instructions in Hinglish or Hindi. These task instructions will be given to a GPT model and we will evaluate the GPT model for completing the instructions. Here are the requirements: 1. Try not to repeat the verb for each instruction to maximize diversity. 2. The language used for the instruction also should be diverse. For example… See the full description on the dataset page: https://huggingface.co/datasets/smangrul/hinglish_self_instruct_v0.texttext-generation1K<n<10K8 likes41 downloads3y agoHugging Face05HachiML /self-rewarding_instruct self-rewarding_instruct self-rewarding_instructは、 Stanford Alpacaの手法 kunishou/oasst1-89k-jaをseed tasksとして mistralai/Mixtral-8x22B-Instruct-v0.1 で作った合成データ(Synthetic data)です。モデルの利用にはDeepinfraを利用しています。データセットはmistralai/Mixtral-8x22B-Instruct-v0.1によって精査されています。 self-rewarding用に作成したため、output_exampleとなっていますがInstruction Tuningにも用いれると思います。 Dataset Details Dataset Description Curated by: HachiMLLanguage(s) (NLP): Japanese License: Apache 2.0 Github:… See the full description on the dataset page: https://huggingface.co/datasets/HachiML/self-rewarding_instruct.texttext-generation10K<n<100K1 likes37 downloads2y agoHugging Face06Aratako /Self-Instruct-Qwen2.5-72B-Instruct-60k Self-Instruct-Qwen2.5-72B-Instruct-60k 概要 以下の手順で作成した約6万件の日本語の合成instructionデータセットです。 MagpieとEvol-Instructを使って作成された合成instructionデータであるAratako/Magpie-Tanuki-Instruction-Selected-Evolved-26.5kに対し、Qwen/Qwen2.5-72B-Instruct-GPTQ-Int8を使ってinstructionのカテゴリを付与 同じカテゴリに分類された3つのinstructionをseed taskとしてQwen/Qwen2.5-72B-Instruct-GPTQ-Int8に与え、Self-Instructの手法で7個のinstructionを生成 元論文の実装とは一部異なります。 generated_instruction列が生成されたinstructionです。 ライセンス 基本的にはApache 2.0に準じますが、Qwen… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Self-Instruct-Qwen2.5-72B-Instruct-60k.texttext-generation10K<n<100K0 likes37 downloads2y agoHugging Face07PJMixers-Dev /bigcode_self-oss-instruct-sc2-exec-filter-50k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT bigcode_self-oss-instruct-sc2-exec-filter-50k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT bigcode/self-oss-instruct-sc2-exec-filter-50k with responses regenerated with gemini-2.0-flash-thinking-exp-1219. Generation Details If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped. If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped. If ["candidates"][0]["finish_reason"] != 1 the sample was… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/bigcode_self-oss-instruct-sc2-exec-filter-50k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.texttext-generation1K<n<10K0 likes31 downloads2y agoHugging Face08iam-ajaymeena /Self-Instruct-Japanese-Elzya-13BA Japanese dataset generated with an opensource elyza/ELYZA-japanese-Llama-2-13b-instruct model. This dataset is used in evaluating AI-generated text detection methods and is well-suited for self-instruct methods. The instructions were taken from: https://huggingface.co/datasets/CausalLM/GPT-4-Self-Instruct-Japanese The model used is: https://huggingface.co/elyza/ELYZA-japanese-Llama-2-7b License: refer to the model's license texttext-generation1K<n<10K0 likes27 downloads2y agoHugging Face09iam-ajaymeena /Self-Instruct-Japanese-Qwen1.5-14BA Japanese dataset generated with Qwen/Qwen1.5-14B model. This dataset is used in evaluating AI-generated text detection methods and is well-suited for self-instruct methods. The instructions were taken from: https://huggingface.co/datasets/CausalLM/GPT-4-Self-Instruct-Japanese The model used is: https://huggingface.co/Qwen/Qwen1.5-14B License: Please refer to the model's license texttext-generation1K<n<10K2 likes18 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.