datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nla-av-responses-llama-70b-layer53details_Nexusflow__Athene-70B
Dataset Card for Evaluation run of Nexusflow/Athene-70B
Dataset automatically created during the evaluation run of model Nexusflow/Athene-70B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Nexusflow__Athene-70B.latenet-v0-activations-llama3.1-70b-base
meta-llama/Llama-3.1-70B — Activation Dataset
Cached activations extracted from meta-llama/Llama-3.1-70B (revision 349b2ddb53ce8f2849a6c168a81980ab25258dac).
Full-sequence activations (80 layers, 8192 dim, float16, all tokens) from meta-llama/Llama-3.1-70B (base) on 23724 LateNet v0 statements (affirmative + negated). Extracted via NDIF. Raw statements only (no chat template). Prompts ordered by negated→generator→pair_id for contiguous domain shards.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/latenet-v0-activations-llama3.1-70b-base.nla-av-ar-attribution-llama-70b-layer53got-activations-llama3.1-70b-base
meta-llama/Llama-3.1-70B — Activation Dataset
Cached activations extracted from meta-llama/Llama-3.1-70B (revision 349b2ddb53ce8f2849a6c168a81980ab25258dac).
Geometry of Truth curated dataset activations for Llama 3.1 70B base
Contents
Tensor
Layers
Dim
Pooling
Shards
Row Bytes
hidden_layers
0-79
8192
-
4
-
Prompts: 7660
Format version: 2.0
Load with lmprobe
from lmprobe import load_activations, Probe
acts =… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/got-activations-llama3.1-70b-base.Magpie-Llama-3.1-70B-Instruct-UnfilteredDataset generated using meta-llama/Llama-3.1-70B-Instruc with the MAGPIE codebase.
The filtered dataset can be found here: HiTZ/Magpie-Llama-3.1-70B-Instruct-Filtered
System prompts used
General
<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nCutting Knowledge Date: December 2023\nToday Date: 26 Jul 2024\n\n<|eot_id|><|start_header_id|>user<|end_header_id|>\n\n
Code
<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are an AI… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3.1-70B-Instruct-Unfiltered.eval-Hermes-4-70B-nonreasoning
hermes-70b-nonreasoning Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.095
math_pass@1:64_samples
64
99.4%
aime25
0.073
math_pass@1:64_samples
64
98.2%
arenahard
0.568
eval/overall_winrate
500
0.0%
bbh_generative
0.805
extractive_match
1
100.0%
creative-writing-v3
0.491
creative_writing_score
96
0.0%
drop_generative_nous
0.784
drop_acc
1
100.0%
eqbench3
0.739
eqbench_score
135
0.0%
gpqa_diamond
0.333… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-70B-nonreasoning.eval-Cogito-v2-preview-70B-reasoning
cogito-thinking Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.322
math_pass@1:64_samples
64
35.2%
aime25
0.221
math_pass@1:64_samples
64
33.3%
arenahard
0.869
eval/overall_winrate
500
0.0%
bbh_generative
0.893
extractive_match
1
2.9%
creative-writing-v3
0.636
creative_writing_score
96
0.0%
drop_generative_nous
0.860
drop_acc
1
0.8%
eqbench3
0.657
eqbench_score
135
0.0%
gpqa_diamond
0.591
gpqa_pass@1:8_samples8… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Cogito-v2-preview-70B-reasoning.eval-Cogito-v2-preview-70B-nonreasoning
cogito-70b-nonthinking Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.122
math_pass@1:64_samples
64
100.0%
aime25
0.060
math_pass@1:64_samples
64
100.0%
arenahard
0.819
eval/overall_winrate
500
0.0%
bbh_generative
0.876
extractive_match
1
100.0%
creative-writing-v3
0.655
creative_writing_score
96
0.0%
drop_generative_nous
0.841
drop_acc
1
100.0%
eqbench3
0.681
eqbench_score
135
0.0%
gpqa_diamond
0.528… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Cogito-v2-preview-70B-nonreasoning.details_sambanovasystems__SambaLingo-Arabic-Chat-70B
Dataset Card for Evaluation run of sambanovasystems/SambaLingo-Arabic-Chat-70B
Dataset automatically created during the evaluation run of model sambanovasystems/SambaLingo-Arabic-Chat-70B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_sambanovasystems__SambaLingo-Arabic-Chat-70B.details_MaziyarPanahi__calme-2.3-llama3-70b
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.3-llama3-70b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.3-llama3-70b.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_MaziyarPanahi__calme-2.3-llama3-70b.Magpie-Llama-3-70B-Instruct-UnfilteredDataset generated using meta-llama/Meta-Llama-3-70B-Instruct with the MAGPIE codebase.
The filtered dataset can be found here: HiTZ/Magpie-Llama-3-70B-Instruct-Filtered
System prompts used
General
<|begin_of_text|><|start_header_id|>user<|end_header_id|>\n\n
Code
<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are an AI assistant designed to provide helpful, step-by-step guidance on coding problems. The user will ask you a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3-70B-Instruct-Unfiltered.Llama-3.3-70B-Instruct-eval-logs-and-scoresreasoning-multilingual-R1-Llama-70B-train
lightblue/reasoning-multilingual-R1-Llama-70B-train
This is a multilingual reasoning dataset covering more than 30 languages.
This dataset was made by:
Sampling prompts from English datasets and translating them to various languages
Generating responses to these prompts 8 times using deepseek-ai/DeepSeek-R1-Distill-Llama-70B
Filtering out <think> sections with incorrect language, non-fluent language, and incorrect answers
This dataset was then used to train a multilingual… See the full description on the dataset page: https://huggingface.co/datasets/lightblue/reasoning-multilingual-R1-Llama-70B-train.Llama-3-Taiwan-70B-Instruct-eval-logs-and-scoreseval-Hermes-4-70B-reasoning
hermes-4-70b-reasoning-40k Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.735
math_pass@1:64_samples
64
8.4%
aime25
0.674
math_pass@1:64_samples
64
9.6%
arenahard
0.901
eval/overall_winrate
500
0.0%
bbh_generative
0.878
extractive_match
1
4.8%
creative-writing-v3
0.775
creative_writing_score
96
0.0%
drop_generative_nous
0.850
drop_acc
1
1.4%
eqbench3
0.847
eqbench_score
135
0.0%
gpqa_diamond
0.661… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-70B-reasoning.details_airev-ai__Amal-70b-v5
Dataset Card for Evaluation run of airev-ai/Amal-70b-v5
Dataset automatically created during the evaluation run of model airev-ai/Amal-70b-v5.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_airev-ai__Amal-70b-v5.details_meta-llama__Llama-2-70b-hf
Dataset Card for Evaluation run of meta-llama/Llama-2-70b-hf
Dataset automatically created during the evaluation run of model meta-llama/Llama-2-70b-hf.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_meta-llama__Llama-2-70b-hf.Magpie-Llama-3-70B-Instruct-FilteredDataset generated using meta-llama/Meta-Llama-3-70B-Instruct with the MAGPIE codebase.
The unfiltered dataset can be found here: HiTZ/Magpie-Llama-3-70B-Instruct-Unfiltered
Filter criteria
def high_quality_filter(example):
return (
example["input_quality"] in ["good", "excellent", "average"]
and example["instruct_reward"] > -10
and not example["instruction"].endswith(":")
and (
example["min_similar_conversation_id"] is None… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3-70B-Instruct-Filtered.details_deepseek-ai__DeepSeek-R1-Distill-Llama-70B
Dataset Card for Evaluation run of deepseek-ai/DeepSeek-R1-Distill-Llama-70B
Dataset automatically created during the evaluation run of model deepseek-ai/DeepSeek-R1-Distill-Llama-70B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_deepseek-ai__DeepSeek-R1-Distill-Llama-70B.details_meta-llama__Meta-Llama-3.1-70B
Dataset Card for Evaluation run of meta-llama/Meta-Llama-3.1-70B
Dataset automatically created during the evaluation run of model meta-llama/Meta-Llama-3.1-70B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_meta-llama__Meta-Llama-3.1-70B.details_meta-llama__Meta-Llama-3-70B-Instruct
Dataset Card for Evaluation run of meta-llama/Meta-Llama-3-70B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Meta-Llama-3-70B-Instruct.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_meta-llama__Meta-Llama-3-70B-Instruct.details_MaziyarPanahi__calme-2.3-llama3.1-70b
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.3-llama3.1-70b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.3-llama3.1-70b.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_MaziyarPanahi__calme-2.3-llama3.1-70b.details_MaziyarPanahi__calme-2.4-llama3-70b
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.4-llama3-70b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.4-llama3-70b.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_MaziyarPanahi__calme-2.4-llama3-70b.llamatales-gre-70b70B_normal_llama_33_70b_instruct__swe_bench_verified_mini
Inspect Dataset: 70B_normal_llama_33_70b_instruct__swe_bench_verified_mini
Dataset Information
This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-17.
Model Information
Model: vllm/meta-llama/Llama-3.3-70B-Instruct
Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'tool_call_parser': 'llama3_json', 'enable_auto_tool_choice': '', 'chat_template':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/70B_normal_llama_33_70b_instruct__swe_bench_verified_mini.Magpie-Llama-3.1-70B-Instruct-FilteredDataset generated using meta-llama/Llama-3.1-70B-Instruct with the MAGPIE codebase.
The unfiltered dataset can be found here: /HiTZ/Magpie-Llama-3.1-70B-Instruct-Unfiltered
Filter criteria
min_repetition = 100
def test_no_repetition(text: str):
# Count the frequency of each word in the text
word_count = Counter(text.split())
# Check if any word appears more than min_repetition times
return all(count <= min_repetition for count in word_count.values())
def… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3.1-70B-Instruct-Filtered.justicio-BOE-A-1978-31229-constitucion-by-articles-qa-qa-groq_llama3_70b_8192-sas
Dataset summary
It is an end-to-end evaluation dataset (using SAS metric) for Justicio.
Domain: Legal, Law, Spanish Constitution
Language: Spanish
SAS summary
The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth.
Justicio summary
Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-qa-groq_llama3_70b_8192-sas.meta-llama_Llama-3.1-70B-Instruct_cot_mmlu-pro_20241017_082906
Dataset Card for meta-llama_Llama-3.1-70B-Instruct_cot_mmlu-pro_20241017_082906
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dvilasuero/meta-llama_Llama-3.1-70B-Instruct_cot_mmlu-pro_20241017_082906/raw/main/pipeline.yaml"
or explore the… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/meta-llama_Llama-3.1-70B-Instruct_cot_mmlu-pro_20241017_082906.justicio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas
Dataset summary
It is an end-to-end evaluation dataset (using SAS metric) for Justicio.
Domain: Legal, Law, Spanish Constitution
Language: Spanish
SAS summary
The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth.
Justicio summary
Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas.
