CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01weizhiwang /Open-Qwen2VL-Data Introduction This repository contains the data for Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources. Project page: https://victorwz.github.io/Open-Qwen2VL Code: https://github.com/Victorwz/Open-Qwen2VL Dataset ccs_ebdataset: CC3M-CC12M-SBU filtered by CLIP, we directly download the webdataset based on the released of curated subset of BLIP-1 datacomp_medium_dfn_webdataset: DataComp-Medium-128M filtered by DFN, we… See the full description on the dataset page: https://huggingface.co/datasets/weizhiwang/Open-Qwen2VL-Data.imageimage-text-to-text10M<n<100M25 likes13k downloads1y agoHugging Face02Magpie-Align /Magpie-Qwen2.5-Pro-300K-Filtered Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Qwen2.5-Pro-300K-Filtered.tabular100K<n<1M14 likes2.9k downloads2y agoHugging Face03Magpie-Align /Magpie-Qwen2.5-Pro-1M-v0.1 Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Qwen2.5-Pro-1M-v0.1.tabulartext-generation1M<n<10M19 likes2.2k downloads2y agoHugging Face04SaylorTwift /RULER-8192-Qwen2.5-3B-tokenizertabular1K<n<10K0 likes1.6k downloads1y agoHugging Face05ENSEONG /full-math-private-n256-Qwen2.5-3B-Instruct-bontabular100K<n<1M0 likes1.5k downloads6mo agoHugging Face06Magpie-Align /Magpie-Qwen2-Pro-200K-Chinese Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Qwen2-Pro-200K-Chinese.tabularquestion-answering100K<n<1M84 likes1.3k downloads2y agoHugging Face07ENSEONG /full-math-private-Qwen2.5-3B-Instruct-bontabular100K<n<1M0 likes772 downloads6mo agoHugging Face08Magpie-Align /Magpie-Qwen2.5-Math-Pro-300K-v0.1 Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Qwen2.5-Math-Pro-300K-v0.1.tabular100K<n<1M2 likes734 downloads2y agoHugging Face09ayeshag7 /fineweb-edu-2019-qwen2 FineWeb-Edu 2019, Qwen2-7B token counts 99,870,012 documents, 99,999,986,613 Qwen2-7B tokens, prepared for continued pretraining as part of the FinMoE project. column type meaning date int32 the FineWeb year, 2019 text string document text, unmodified token_count int32 Qwen2-7B tokens in text Source HuggingFaceFW/fineweb-edu at revision 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9, all 12 CommonCrawl dumps of 2019 (190 shards). Token… See the full description on the dataset page: https://huggingface.co/datasets/ayeshag7/fineweb-edu-2019-qwen2.tabular10M<n<100M0 likes694 downloads12d agoHugging Face10ENSEONG /stratified-solvable-1k-math-private-Qwen2.5-3B-Instruct-bontabular10K<n<100K0 likes687 downloads6mo agoHugging Face11ayeshag7 /fineweb-edu-2015-qwen2 FineWeb-Edu 2015, Qwen2-7B token counts 93,077,934 documents, 99,999,999,500 Qwen2-7B tokens, prepared for continued pretraining as part of the FinMoE project. column type meaning date int32 the FineWeb year, 2015 text string document text, unmodified token_count int32 Qwen2-7B tokens in text Source HuggingFaceFW/fineweb-edu at revision 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9, all 10 CommonCrawl dumps of 2015 (134 shards). Token… See the full description on the dataset page: https://huggingface.co/datasets/ayeshag7/fineweb-edu-2015-qwen2.tabular10M<n<100M0 likes630 downloads12d agoHugging Face12Lansechen /details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval. The dataset is composed of 5 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 23 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval.tabular1K<n<10K0 likes613 downloads1y agoHugging Face13latent-lab /got-activations-qwen2.5-0.5b Qwen/Qwen2.5-0.5B — Activation Dataset Cached activations extracted from Qwen/Qwen2.5-0.5B (revision 060db6499f32faf8b98477b0a26969ef7d8b9987). Full-sequence activations (24 layers, 896 dim, float16) and top-100 logits from Qwen/Qwen2.5-0.5B on 7,660 Geometry of Truth statements. Per-layer sharding (v1.2) with independent shard boundaries. Contents Tensor Layers Dim Pooling Shards Row Bytes hidden_layers 0-23 896 - 1 - logits_topk - k=100 last_token 1 1200… See the full description on the dataset page: https://huggingface.co/datasets/latent-lab/got-activations-qwen2.5-0.5b.tabularfeature-extraction1K<n<10K0 likes573 downloads6mo agoHugging Face14stevenyuan666 /fineweb-edu-2013-qwen2-7b FineWeb-Edu 2013 with Qwen2-7B token counts Every 2013 FineWeb-Edu document, prepared for continued pretraining, with token counts computed by a pinned Qwen2-7B tokenizer. The pipeline is year-agnostic: the year, source revision, tokenizer contract, and selection rule all come from a config file. 2013 uses processing_config.json. The 2017 companion dataset, which is large enough to require shuffling and a token budget rather than retaining everything, is at… See the full description on the dataset page: https://huggingface.co/datasets/stevenyuan666/fineweb-edu-2013-qwen2-7b.tabulartext-generation10M<n<100M0 likes543 downloads13d agoHugging Face15OALL /details_MaziyarPanahi__calme-2.7-qwen2-7b Dataset Card for Evaluation run of MaziyarPanahi/calme-2.7-qwen2-7b Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.7-qwen2-7b. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_MaziyarPanahi__calme-2.7-qwen2-7b.tabular100K<n<1M0 likes518 downloads2y agoHugging Face16SaylorTwift /RULER-32768-Qwen2.5-3B-tokenizertabular1K<n<10K0 likes489 downloads1y agoHugging Face17Lansechen /details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-default Dataset Card for Evaluation run of Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-default Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-default. The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 11 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-default.tabular1K<n<10K0 likes483 downloads1y agoHugging Face18mlfoundations-dev /Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179 mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 60.7 90.8 89.4 63.2 52.4 48.5 27.4 26.2 48.3 12.0 34.3 34.7 AIME24 Average Accuracy: 60.67% ± 2.25% Number of Runs: 10 Run Accuracy Questions Solved Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179.tabular10K<n<100K1 likes475 downloads1y agoHugging Face19mlfoundations-dev /Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_2870 mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_2870 Precomputed model outputs for evaluation. Evaluation Results AIME24 Average Accuracy: 60.67% ± 2.20% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 70.00% 21 30 2 53.33% 16 30 3 53.33% 16 30 4 66.67% 20 30 5 63.33% 19 30 6 66.67% 20 30 7 60.00% 18 30 8 46.67% 14 30 9 63.33% 19 30 10 63.33% 19 30 tabularn<1K0 likes451 downloads1y agoHugging Face20hanlincs /in1k_clip_qwen25vl_3b_224res_64tokens_new_pttabular1M<n<10M0 likes423 downloads1y agoHugging Face21hanlincs /in1k_clip_qwen25vl_3b_448res_256tokens_new_merged_pttabular1M<n<10M0 likes393 downloads1y agoHugging Face22Magpie-Align /Magpie-Qwen2-Pro-200K-English Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Qwen2-Pro-200K-English.tabular100K<n<1M9 likes380 downloads2y agoHugging Face23Usmansafder /fineweb-edu-2018-qwen2 FineWeb-Edu 2018, Qwen2-7B token counts 103,794,847 documents, 99,999,999,271 Qwen2-7B tokens, prepared for continued pretraining as part of the FinMoE project. column type meaning date int32 the FineWeb year, 2018 text string document text, unmodified token_count int32 Qwen2-7B tokens in text Source HuggingFaceFW/fineweb-edu at revision 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9, all 12 CommonCrawl dumps of 2018 (196 shards). Token… See the full description on the dataset page: https://huggingface.co/datasets/Usmansafder/fineweb-edu-2018-qwen2.tabular100M<n<1B0 likes366 downloads11d agoHugging Face24gist-sparse-attention /GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk4-chunk4. Each sample is tokenized and formatted with GSA gist tokens for continue pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4 — model trained on this dataset tabular10K<n<100K0 likes361 downloads6mo agoHugging Face25mlfoundations-dev /Qwen2.5-1.5B-Instruct_eval_5554 mlfoundations-dev/Qwen2.5-1.5B-Instruct_eval_5554 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces HLE HMMT AIME25 LiveCodeBenchv5 Accuracy 3.0 30.8 50.2 32.5 16.4 24.7 5.5 0.8 2.2 15.3 0.0 0.7 5.1 AIME24 Average Accuracy: 3.00% ± 0.88% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 3.33% 1… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Qwen2.5-1.5B-Instruct_eval_5554.tabular10K<n<100K0 likes352 downloads1y agoHugging Face26Usmansafder /fineweb-edu-2014-qwen2 FineWeb-Edu 2014, Qwen2-7B token counts 86,901,732 documents, 92,348,618,822 Qwen2-7B tokens, prepared for continued pretraining as part of the FinMoE project. column type meaning date int32 the FineWeb year, 2014 text string document text, unmodified token_count int32 Qwen2-7B tokens in text Source HuggingFaceFW/fineweb-edu at revision 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9, all 8 CommonCrawl dumps of 2014 (115 shards). Token… See the full description on the dataset page: https://huggingface.co/datasets/Usmansafder/fineweb-edu-2014-qwen2.tabular10M<n<100M0 likes336 downloads11d agoHugging Face27OALL /details_deep-analysis-research__D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2_v2 Dataset Card for Evaluation run of deep-analysis-research/D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2 Dataset automatically created during the evaluation run of model deep-analysis-research/D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2. The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_deep-analysis-research__D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2_v2.tabular100K<n<1M0 likes302 downloads1y agoHugging Face28deep-analysis-research /details_D2IL-Arabic-Qwen2.5-72B-Instruct-v0.1tabular10K<n<100K0 likes299 downloads1y agoHugging Face29Magpie-Align /Magpie-Qwen2-Air-3M-v0.1 Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Qwen2-Air-3M-v0.1.tabular1M<n<10M6 likes277 downloads2y agoHugging Face30ENSEONG /full-gsm8k-private-n256-Qwen2.5-3B-Instruct-bontabular10K<n<100K0 likes263 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.