datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Open-Qwen2VL-Data
Introduction
This repository contains the data for Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources.
Project page: https://victorwz.github.io/Open-Qwen2VL
Code: https://github.com/Victorwz/Open-Qwen2VL
Dataset
ccs_ebdataset: CC3M-CC12M-SBU filtered by CLIP, we directly download the webdataset based on the released of curated subset of BLIP-1
datacomp_medium_dfn_webdataset: DataComp-Medium-128M filtered by DFN, we… See the full description on the dataset page: https://huggingface.co/datasets/weizhiwang/Open-Qwen2VL-Data.Magpie-Qwen2.5-Pro-300K-Filtered
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Qwen2.5-Pro-300K-Filtered.Magpie-Qwen2.5-Pro-1M-v0.1
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Qwen2.5-Pro-1M-v0.1.RULER-8192-Qwen2.5-3B-tokenizerfull-math-private-n256-Qwen2.5-3B-Instruct-bonMagpie-Qwen2-Pro-200K-Chinese
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Qwen2-Pro-200K-Chinese.full-math-private-Qwen2.5-3B-Instruct-bonMagpie-Qwen2.5-Math-Pro-300K-v0.1
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Qwen2.5-Math-Pro-300K-v0.1.fineweb-edu-2019-qwen2
FineWeb-Edu 2019, Qwen2-7B token counts
99,870,012 documents, 99,999,986,613 Qwen2-7B tokens, prepared for continued pretraining as part of the
FinMoE project.
column
type
meaning
date
int32
the FineWeb year, 2019
text
string
document text, unmodified
token_count
int32
Qwen2-7B tokens in text
Source
HuggingFaceFW/fineweb-edu at revision 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9, all 12
CommonCrawl dumps of 2019 (190 shards).
Token… See the full description on the dataset page: https://huggingface.co/datasets/ayeshag7/fineweb-edu-2019-qwen2.stratified-solvable-1k-math-private-Qwen2.5-3B-Instruct-bonfineweb-edu-2015-qwen2
FineWeb-Edu 2015, Qwen2-7B token counts
93,077,934 documents, 99,999,999,500 Qwen2-7B tokens, prepared for continued pretraining as part of the
FinMoE project.
column
type
meaning
date
int32
the FineWeb year, 2015
text
string
document text, unmodified
token_count
int32
Qwen2-7B tokens in text
Source
HuggingFaceFW/fineweb-edu at revision 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9, all 10
CommonCrawl dumps of 2015 (134 shards).
Token… See the full description on the dataset page: https://huggingface.co/datasets/ayeshag7/fineweb-edu-2015-qwen2.details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval.
The dataset is composed of 5 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 23 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval.got-activations-qwen2.5-0.5b
Qwen/Qwen2.5-0.5B — Activation Dataset
Cached activations extracted from Qwen/Qwen2.5-0.5B (revision 060db6499f32faf8b98477b0a26969ef7d8b9987).
Full-sequence activations (24 layers, 896 dim, float16) and top-100 logits from Qwen/Qwen2.5-0.5B on 7,660 Geometry of Truth statements. Per-layer sharding (v1.2) with independent shard boundaries.
Contents
Tensor
Layers
Dim
Pooling
Shards
Row Bytes
hidden_layers
0-23
896
-
1
-
logits_topk
-
k=100
last_token
1
1200… See the full description on the dataset page: https://huggingface.co/datasets/latent-lab/got-activations-qwen2.5-0.5b.fineweb-edu-2013-qwen2-7b
FineWeb-Edu 2013 with Qwen2-7B token counts
Every 2013 FineWeb-Edu document, prepared for continued pretraining, with token
counts computed by a pinned Qwen2-7B tokenizer.
The pipeline is year-agnostic: the year, source revision, tokenizer contract,
and selection rule all come from a config file. 2013 uses
processing_config.json. The 2017 companion dataset, which is large enough to
require shuffling and a token budget rather than retaining everything, is at… See the full description on the dataset page: https://huggingface.co/datasets/stevenyuan666/fineweb-edu-2013-qwen2-7b.details_MaziyarPanahi__calme-2.7-qwen2-7b
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.7-qwen2-7b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.7-qwen2-7b.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_MaziyarPanahi__calme-2.7-qwen2-7b.RULER-32768-Qwen2.5-3B-tokenizerdetails_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-default
Dataset Card for Evaluation run of Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-default
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-default.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 11 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-default.Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179
mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
60.7
90.8
89.4
63.2
52.4
48.5
27.4
26.2
48.3
12.0
34.3
34.7
AIME24
Average Accuracy: 60.67% ± 2.25%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179.Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_2870
mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_2870
Precomputed model outputs for evaluation.
Evaluation Results
AIME24
Average Accuracy: 60.67% ± 2.20%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
70.00%
21
30
2
53.33%
16
30
3
53.33%
16
30
4
66.67%
20
30
5
63.33%
19
30
6
66.67%
20
30
7
60.00%
18
30
8
46.67%
14
30
9
63.33%
19
30
10
63.33%
19
30
in1k_clip_qwen25vl_3b_224res_64tokens_new_ptin1k_clip_qwen25vl_3b_448res_256tokens_new_merged_ptMagpie-Qwen2-Pro-200K-English
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Qwen2-Pro-200K-English.fineweb-edu-2018-qwen2
FineWeb-Edu 2018, Qwen2-7B token counts
103,794,847 documents, 99,999,999,271 Qwen2-7B tokens, prepared for continued pretraining as part of the
FinMoE project.
column
type
meaning
date
int32
the FineWeb year, 2018
text
string
document text, unmodified
token_count
int32
Qwen2-7B tokens in text
Source
HuggingFaceFW/fineweb-edu at revision 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9, all 12
CommonCrawl dumps of 2018 (196 shards).
Token… See the full description on the dataset page: https://huggingface.co/datasets/Usmansafder/fineweb-edu-2018-qwen2.GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data
GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk4-chunk4.
Each sample is tokenized and formatted with GSA gist tokens for continue pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4 — model trained on this dataset
Qwen2.5-1.5B-Instruct_eval_5554
mlfoundations-dev/Qwen2.5-1.5B-Instruct_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
HLE
HMMT
AIME25
LiveCodeBenchv5
Accuracy
3.0
30.8
50.2
32.5
16.4
24.7
5.5
0.8
2.2
15.3
0.0
0.7
5.1
AIME24
Average Accuracy: 3.00% ± 0.88%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
3.33%
1… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Qwen2.5-1.5B-Instruct_eval_5554.fineweb-edu-2014-qwen2
FineWeb-Edu 2014, Qwen2-7B token counts
86,901,732 documents, 92,348,618,822 Qwen2-7B tokens, prepared for continued pretraining as part of the
FinMoE project.
column
type
meaning
date
int32
the FineWeb year, 2014
text
string
document text, unmodified
token_count
int32
Qwen2-7B tokens in text
Source
HuggingFaceFW/fineweb-edu at revision 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9, all 8
CommonCrawl dumps of 2014 (115 shards).
Token… See the full description on the dataset page: https://huggingface.co/datasets/Usmansafder/fineweb-edu-2014-qwen2.details_deep-analysis-research__D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2_v2
Dataset Card for Evaluation run of deep-analysis-research/D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2
Dataset automatically created during the evaluation run of model deep-analysis-research/D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2.
The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_deep-analysis-research__D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2_v2.details_D2IL-Arabic-Qwen2.5-72B-Instruct-v0.1Magpie-Qwen2-Air-3M-v0.1
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Qwen2-Air-3M-v0.1.full-gsm8k-private-n256-Qwen2.5-3B-Instruct-bon
