datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
repro-organic-data-72BTaur_CoT_Analysis_Project___Qwen__Qwen2-72B-Instructnvidia_NVLM-D-72B-jdgfct-Readabilitydetails_deep-analysis-research__D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2_v2
Dataset Card for Evaluation run of deep-analysis-research/D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2
Dataset automatically created during the evaluation run of model deep-analysis-research/D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2.
The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_deep-analysis-research__D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2_v2.nvidia_NVLM-D-72B-jdgfct-Completenessrepro-rephrased-data-72BThis is the 72B rephrased data by repro-rephraser-4B from RePro: Training Language Models to Faithfully Recycle the Web for Pretraining.
Code: https://github.com/cxcscmu/RePro
details_AbdulmalekDS__qwen72b-ar-lora_v2
Dataset Card for Evaluation run of AbdulmalekDS/qwen72b-ar-lora
Dataset automatically created during the evaluation run of model AbdulmalekDS/qwen72b-ar-lora.
The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_AbdulmalekDS__qwen72b-ar-lora_v2.details_tunny__Arabic_Qwen2.5_72B_instruct_finetune_0.1_v2
Dataset Card for Evaluation run of tunny/Arabic_Qwen2.5_72B_instruct_finetune_0.1
Dataset automatically created during the evaluation run of model tunny/Arabic_Qwen2.5_72B_instruct_finetune_0.1.
The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_tunny__Arabic_Qwen2.5_72B_instruct_finetune_0.1_v2.details_D2IL-Arabic-Qwen2.5-72B-Instruct-v0.1details_Qwen__Qwen1.5-72B
Dataset Card for Evaluation run of Qwen/Qwen1.5-72B
Dataset automatically created during the evaluation run of model Qwen/Qwen1.5-72B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen1.5-72B.details_MaziyarPanahi__calme-2.3-qwen2-72b
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.3-qwen2-72b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.3-qwen2-72b.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_MaziyarPanahi__calme-2.3-qwen2-72b.allenai_WildChat-1M-Full-Qwen_Qwen2.5-72B-Instruct-lcdetails_deep-analysis-research__D2IL-Arabic-Qwen2.5-72B-Instruct-v0.1_v2
Dataset Card for Evaluation run of deep-analysis-research/D2IL-Arabic-Qwen2.5-72B-Instruct-v0.1
Dataset automatically created during the evaluation run of model deep-analysis-research/D2IL-Arabic-Qwen2.5-72B-Instruct-v0.1.
The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_deep-analysis-research__D2IL-Arabic-Qwen2.5-72B-Instruct-v0.1_v2.details_freewheelin__free-evo-qwen72b-v0.8-re
Dataset Card for Evaluation run of freewheelin/free-evo-qwen72b-v0.8-re
Dataset automatically created during the evaluation run of model freewheelin/free-evo-qwen72b-v0.8-re.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_freewheelin__free-evo-qwen72b-v0.8-re.xlam-irrelevance-7.5k-qwen2.5-72b-distill-parsed
[PARSED] xlam-irrelevance-7.5k
The data in this dataset is a version of the original MadeAgents/xlam-irrelevance-7.5k response generated with Qwen2.5 72B. It provides the remainder with some cases where the 72B model chose to call a function during generation excluded.
Overview
The xlam-irrelevance-7.5k is a specialized dataset designed to activate the ability of irrelevant function detection for large language models (LLMs).
Source and Construction
This… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/xlam-irrelevance-7.5k-qwen2.5-72b-distill-parsed.all-Qwen2.5-72B-Instruct-AWQdetails_Replete-AI__Replete-LLM-V2.5-Qwen-72b
Dataset Card for Evaluation run of Replete-AI/Replete-LLM-V2.5-Qwen-72b
Dataset automatically created during the evaluation run of model Replete-AI/Replete-LLM-V2.5-Qwen-72b.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Replete-AI__Replete-LLM-V2.5-Qwen-72b.details_Qwen__Qwen2.5-Math-72B
Dataset Card for Evaluation run of Qwen/Qwen2.5-Math-72B
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Math-72B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen2.5-Math-72B.Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k
Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k
概要
5種類のオープンモデルとQwen/Qwen2.5-72B-Instruct-GPTQ-Int8を使って作成した、190854件の日本語合成Preferenceデータセットです。
以下、データセットの詳細です。
instructionには、Aratako/Magpie-Tanuki-8B-annotated-96kのinput_qualityがexcellentのものを利用
回答生成には、以下の5つのApache 2.0ライセンスのモデルを利用
weblab-GENIAC/Tanuki-8B-dpo-v1.0
team-hatakeyama-phase2/Tanuki-8x8B-dpo-v1.0-GPTQ-8bit
cyberagent/calm3-22b-chat
llm-jp/llm-jp-3-13b-instruct
Qwen/Qwen2.5-32B-Instruct-GPTQ-Int8… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k.xlam-irrelevance-7.5k-qwen2.5-72b-distill-hermesdetails_davidkim205__Rhea-72b-v0.5
Dataset Card for Evaluation run of davidkim205/Rhea-72b-v0.5
Dataset automatically created during the evaluation run of model davidkim205/Rhea-72b-v0.5.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_davidkim205__Rhea-72b-v0.5.details_Qwen__Qwen2.5-72B-Instruct
Dataset Card for Evaluation run of Qwen/Qwen2.5-72B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-72B-Instruct.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen2.5-72B-Instruct.Qwen__Qwen2.5-72B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-72B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-72B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-72B-Instruct-details.details_Qwen__Qwen2.5-72B
Dataset Card for Evaluation run of Qwen/Qwen2.5-72B
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-72B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen2.5-72B.details_abacusai__Smaug-72B-v0.1
Dataset Card for Evaluation run of abacusai/Smaug-72B-v0.1
Dataset automatically created during the evaluation run of model abacusai/Smaug-72B-v0.1.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_abacusai__Smaug-72B-v0.1.abacusai__Dracarys-72B-Instruct-details
Dataset Card for Evaluation run of abacusai/Dracarys-72B-Instruct
Dataset automatically created during the evaluation run of model abacusai/Dracarys-72B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Dracarys-72B-Instruct-details.rubenroy__Gilgamesh-72B-details
Dataset Card for Evaluation run of rubenroy/Gilgamesh-72B
Dataset automatically created during the evaluation run of model rubenroy/Gilgamesh-72B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/rubenroy__Gilgamesh-72B-details.qwen-2.5-72b-instruct-eagle-numbers-run-0llm-complex-reasoning-train-qwen2-72b-instruct-correct
Note
Data Seed from 基于封闭世界假设的复杂逻辑推理
Generate from Qwen2-72B-Instruct with prompt
train.jsonl for 推理答案和题目答案一致, no_train.jsonl推理答案和题目答案不一致
注: 题目答案不一定正确
EVA-UNIT-01__EVA-Qwen2.5-72B-v0.2-details
Dataset Card for Evaluation run of EVA-UNIT-01/EVA-Qwen2.5-72B-v0.2
Dataset automatically created during the evaluation run of model EVA-UNIT-01/EVA-Qwen2.5-72B-v0.2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EVA-UNIT-01__EVA-Qwen2.5-72B-v0.2-details.
