datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nla-av-responses-llama-70b-layer53details_Nexusflow__Athene-70B
Dataset Card for Evaluation run of Nexusflow/Athene-70B
Dataset automatically created during the evaluation run of model Nexusflow/Athene-70B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Nexusflow__Athene-70B.latenet-v0-activations-llama3.1-70b-base
meta-llama/Llama-3.1-70B — Activation Dataset
Cached activations extracted from meta-llama/Llama-3.1-70B (revision 349b2ddb53ce8f2849a6c168a81980ab25258dac).
Full-sequence activations (80 layers, 8192 dim, float16, all tokens) from meta-llama/Llama-3.1-70B (base) on 23724 LateNet v0 statements (affirmative + negated). Extracted via NDIF. Raw statements only (no chat template). Prompts ordered by negated→generator→pair_id for contiguous domain shards.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/latenet-v0-activations-llama3.1-70b-base.3112_llm_70b_trainingMagpie-Reasoning-V2-250K-CoT-Deepseek-R1-Llama-70B
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Reasoning-V2-250K-CoT-Deepseek-R1-Llama-70B.nla-av-ar-attribution-llama-70b-layer53got-activations-llama3.1-70b-base
meta-llama/Llama-3.1-70B — Activation Dataset
Cached activations extracted from meta-llama/Llama-3.1-70B (revision 349b2ddb53ce8f2849a6c168a81980ab25258dac).
Geometry of Truth curated dataset activations for Llama 3.1 70B base
Contents
Tensor
Layers
Dim
Pooling
Shards
Row Bytes
hidden_layers
0-79
8192
-
4
-
Prompts: 7660
Format version: 2.0
Load with lmprobe
from lmprobe import load_activations, Probe
acts =… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/got-activations-llama3.1-70b-base.Magpie-Llama-3.1-70B-Instruct-UnfilteredDataset generated using meta-llama/Llama-3.1-70B-Instruc with the MAGPIE codebase.
The filtered dataset can be found here: HiTZ/Magpie-Llama-3.1-70B-Instruct-Filtered
System prompts used
General
<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nCutting Knowledge Date: December 2023\nToday Date: 26 Jul 2024\n\n<|eot_id|><|start_header_id|>user<|end_header_id|>\n\n
Code
<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are an AI… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3.1-70B-Instruct-Unfiltered.meta-llama_Llama-3.1-70B-Instruct-jdgfct-Readabilityeval-Hermes-4-70B-nonreasoning
hermes-70b-nonreasoning Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.095
math_pass@1:64_samples
64
99.4%
aime25
0.073
math_pass@1:64_samples
64
98.2%
arenahard
0.568
eval/overall_winrate
500
0.0%
bbh_generative
0.805
extractive_match
1
100.0%
creative-writing-v3
0.491
creative_writing_score
96
0.0%
drop_generative_nous
0.784
drop_acc
1
100.0%
eqbench3
0.739
eqbench_score
135
0.0%
gpqa_diamond
0.333… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-70B-nonreasoning.eval-Cogito-v2-preview-70B-reasoning
cogito-thinking Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.322
math_pass@1:64_samples
64
35.2%
aime25
0.221
math_pass@1:64_samples
64
33.3%
arenahard
0.869
eval/overall_winrate
500
0.0%
bbh_generative
0.893
extractive_match
1
2.9%
creative-writing-v3
0.636
creative_writing_score
96
0.0%
drop_generative_nous
0.860
drop_acc
1
0.8%
eqbench3
0.657
eqbench_score
135
0.0%
gpqa_diamond
0.591
gpqa_pass@1:8_samples8… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Cogito-v2-preview-70B-reasoning.Nexusflow_Athene-70B-jdgfct-Completenesseval-Cogito-v2-preview-70B-nonreasoning
cogito-70b-nonthinking Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.122
math_pass@1:64_samples
64
100.0%
aime25
0.060
math_pass@1:64_samples
64
100.0%
arenahard
0.819
eval/overall_winrate
500
0.0%
bbh_generative
0.876
extractive_match
1
100.0%
creative-writing-v3
0.655
creative_writing_score
96
0.0%
drop_generative_nous
0.841
drop_acc
1
100.0%
eqbench3
0.681
eqbench_score
135
0.0%
gpqa_diamond
0.528… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Cogito-v2-preview-70B-nonreasoning.Nexusflow_Athene-70B-jdgfct-Harmlessnessdetails_sambanovasystems__SambaLingo-Arabic-Chat-70B
Dataset Card for Evaluation run of sambanovasystems/SambaLingo-Arabic-Chat-70B
Dataset automatically created during the evaluation run of model sambanovasystems/SambaLingo-Arabic-Chat-70B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_sambanovasystems__SambaLingo-Arabic-Chat-70B.Magpie-Reasoning-V2-250K-CoT-Deepseek-R1-Llama-70B-formatteddetails_MaziyarPanahi__calme-2.3-llama3-70b
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.3-llama3-70b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.3-llama3-70b.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_MaziyarPanahi__calme-2.3-llama3-70b.Magpie-Llama-3-70B-Instruct-UnfilteredDataset generated using meta-llama/Meta-Llama-3-70B-Instruct with the MAGPIE codebase.
The filtered dataset can be found here: HiTZ/Magpie-Llama-3-70B-Instruct-Filtered
System prompts used
General
<|begin_of_text|><|start_header_id|>user<|end_header_id|>\n\n
Code
<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are an AI assistant designed to provide helpful, step-by-step guidance on coding problems. The user will ask you a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3-70B-Instruct-Unfiltered.llama-3.1-tulu-3-70b-preference-mixture
Llama 3.1 Tulu 3 70B Preference Mixture
Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact.
This preference mixture used for DPO on our the Llama 3.1 Tulu 3 70B SFT checkpoint to obtain Llama 3.1 Tulu 3 70B DPO.
This mix is made up from the following preference datasets:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/llama-3.1-tulu-3-70b-preference-mixture.eval-Hermes-4-70B-reasoning
hermes-4-70b-reasoning-40k Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.735
math_pass@1:64_samples
64
8.4%
aime25
0.674
math_pass@1:64_samples
64
9.6%
arenahard
0.901
eval/overall_winrate
500
0.0%
bbh_generative
0.878
extractive_match
1
4.8%
creative-writing-v3
0.775
creative_writing_score
96
0.0%
drop_generative_nous
0.850
drop_acc
1
1.4%
eqbench3
0.847
eqbench_score
135
0.0%
gpqa_diamond
0.661… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-70B-reasoning.all-Meta-Llama-3.1-70B-Instruct-AWQ-INT4Llama-3.3-70B-Instruct-eval-logs-and-scoresLlama-3-Taiwan-70B-Instruct-eval-logs-and-scoresreasoning-multilingual-R1-Llama-70B-train
lightblue/reasoning-multilingual-R1-Llama-70B-train
This is a multilingual reasoning dataset covering more than 30 languages.
This dataset was made by:
Sampling prompts from English datasets and translating them to various languages
Generating responses to these prompts 8 times using deepseek-ai/DeepSeek-R1-Distill-Llama-70B
Filtering out <think> sections with incorrect language, non-fluent language, and incorrect answers
This dataset was then used to train a multilingual… See the full description on the dataset page: https://huggingface.co/datasets/lightblue/reasoning-multilingual-R1-Llama-70B-train.details_airev-ai__Amal-70b-v5
Dataset Card for Evaluation run of airev-ai/Amal-70b-v5
Dataset automatically created during the evaluation run of model airev-ai/Amal-70b-v5.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_airev-ai__Amal-70b-v5.GPQA_with_Llama_3.1_70B_Instruct_v1
GPQA with Llama-3.1-70B-Instruct
This dataset contains 646 graduate-level science questions from the GPQA benchmark with 100 candidate responses generated by Llama-3.1-70B-Instruct for each problem. Each response has been evaluated for correctness using a mixture of GPT-4o-mini and procedural Python code to robustly parse different answer formats, and scored by multiple reward models (scalar values) and LM judges (boolean verdicts).
Dataset Structure
Split: Single… See the full description on the dataset page: https://huggingface.co/datasets/hazyresearch/GPQA_with_Llama_3.1_70B_Instruct_v1.details_meta-llama__Llama-2-70b-hf
Dataset Card for Evaluation run of meta-llama/Llama-2-70b-hf
Dataset automatically created during the evaluation run of model meta-llama/Llama-2-70b-hf.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_meta-llama__Llama-2-70b-hf.regolo-instruct-llama70B
Regolo Instruct Llama-3.3-70B - Regolo.ai 🧠
Description
This dataset was generated using Llama-3.3-70B, served via regolo.ai.The generation process was divided into two main stages:
Translation of questions from open-source English-language datasets using Qwen2.5-7B
Response generation through regolo
Data
{
"messages": [
{"role": "system", "content": "<SYSTEM MESSAGE>"},
{"role": "user", "content":… See the full description on the dataset page: https://huggingface.co/datasets/ReDiX/regolo-instruct-llama70B.Magpie-Reasoning-V1-150K-CoT-Deepseek-R1-Llama-70B
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Reasoning-V1-150K-CoT-Deepseek-R1-Llama-70B.Magpie-Llama-3-70B-Instruct-FilteredDataset generated using meta-llama/Meta-Llama-3-70B-Instruct with the MAGPIE codebase.
The unfiltered dataset can be found here: HiTZ/Magpie-Llama-3-70B-Instruct-Unfiltered
Filter criteria
def high_quality_filter(example):
return (
example["input_quality"] in ["good", "excellent", "average"]
and example["instruct_reward"] > -10
and not example["instruction"].endswith(":")
and (
example["min_similar_conversation_id"] is None… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3-70B-Instruct-Filtered.
