CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-llm-leaderboard /Salesforce__LLaMA-3-8B-SFR-Iterative-DPO-R-detailsgated Dataset Card for Evaluation run of Salesforce/LLaMA-3-8B-SFR-Iterative-DPO-R Dataset automatically created during the evaluation run of model Salesforce/LLaMA-3-8B-SFR-Iterative-DPO-R The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Salesforce__LLaMA-3-8B-SFR-Iterative-DPO-R-details.tabular10K<n<100K0 likes56 downloads2y agoHugging Face02Aratako /iterative-dpo-data-for-SimPO-iter2 iterative-dpo-data-for-SimPO-iter2 概要 合成instructionデータであるAratako/Magpie-Tanuki-Instruction-Selected-Evolved-26.5kを元に以下のような手順で作成した日本語Preferenceデータセットです。 開発途中のモデルであるAratako/Llama-Gemma-2-27b-CPO_SimPO-iter1を用いて、temperature=1で回答を5回生成 5個の回答それぞれに対して、Qwen/Qwen2.5-72B-Instruct-GPTQ-Int8を用いて0~5点のスコア付けを実施 1つのinstructionに対する5個の回答について、最もスコアが高いものをchosenに、低いものをrejectedに配置 全て同じスコアの場合や、最も良いスコアが2点以下の場合は除外 ライセンス 本データセットは回答の作成に利用したモデルの関係で以下のライセンスの影響を受けます。 META LLAMA 3.1… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/iterative-dpo-data-for-SimPO-iter2.tabulartext-generation10K<n<100K1 likes41 downloads2y agoHugging Face03Aratako /iterative-dpo-data-for-ORPO-iter3 iterative-dpo-data-for-ORPO-iter3 概要 合成instructionデータであるAratako/Self-Instruct-Qwen2.5-72B-Instruct-60kを元に以下のような手順で作成した日本語Preferenceデータセットです。 開発途中のモデルであるAratako/Llama-Gemma-2-27b-CPO_SimPO-iter2を用いて、temperature=1で回答を5回生成 5個の回答それぞれに対して、Qwen/Qwen2.5-72B-Instruct-GPTQ-Int8を用いて0~5点のスコア付けを実施 1つのinstructionに対する5個の回答について、最もスコアが高いものをchosenに、低いものをrejectedに配置 全て同じスコアの場合や、最も良いスコアが2点以下の場合は除外 ライセンス 本データセットは回答の作成に利用したモデルの関係で以下のライセンスの影響を受けます。 META LLAMA 3.1 COMMUNITY… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/iterative-dpo-data-for-ORPO-iter3.tabulartext-generation10K<n<100K3 likes25 downloads2y agoHugging Face04OALL /details_RLHFlow__LLaMA3-iterative-DPO-final Dataset Card for Evaluation run of RLHFlow/LLaMA3-iterative-DPO-final Dataset automatically created during the evaluation run of model RLHFlow/LLaMA3-iterative-DPO-final. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_RLHFlow__LLaMA3-iterative-DPO-final.tabular100K<n<1M0 likes10 downloads2y agoHugging Face05snap-stanford /preference_iterative_hardtabularn<1K0 likes6 downloads2y agoHugging Face06open-llm-leaderboard /RLHFlow__LLaMA3-iterative-DPO-final-detailsgated Dataset Card for Evaluation run of RLHFlow/LLaMA3-iterative-DPO-final Dataset automatically created during the evaluation run of model RLHFlow/LLaMA3-iterative-DPO-final The dataset is composed of 43 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/RLHFlow__LLaMA3-iterative-DPO-final-details.tabular10K<n<100K0 likes5 downloads2y agoHugging Face07kazuyamaa /iterative-ppo-data-iter3tabular10K<n<100K0 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.