datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SlimPajama-6BSampled version of cerebras/SlimPajama-627B.
Since the original data was shuffled before chunking, I only downloaded train/chunk1 (of 10 total) and further sampled 10%. This should result in roughly 6B tokens, hence SlimPajama-6B.
The dataset is 24GBs in storage size when decompressed (original dataset is over 2TBs) and has 5489000 rows.
The validation set and test set were sampled as well.
Data source proportions for SlimPajama-627B and SlimPajama-6B
For sanity purpose, I… See the full description on the dataset page: https://huggingface.co/datasets/DKYoon/SlimPajama-6B.hacs_segment_internvideo2_6b_w16_s8details_EleutherAI__gpt-j-6b
Dataset Card for Evaluation run of EleutherAI/gpt-j-6b
Dataset Summary
Dataset automatically created during the evaluation run of model EleutherAI/gpt-j-6b on the Open LLM Leaderboard.
The dataset is composed of 122 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 8 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_EleutherAI__gpt-j-6b.SlimPajama-6B_km-ip-d512ClimbMix-6BT
ClimbMix-6BT
This is the tokenized nvidia/Nemotron-ClimbMix (10M subset) using SmolLM2-135M tokenzier. Data is divided into shards (.npy files) for easier to load with PyTorch IterableDataset.
Each .npy file can be loaded with numpy.load('file_name.npy').
Split
# Documents
# Shards
# Tokens
train
9,900,000
65
6,463,974,020 (6.5B)
val
100,000
1
64,859,672 (65M)
Total
10,000,000
66
6,528,833,692 (6.5B)
Example of usage
uvx hf download… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/ClimbMix-6BT.activitynet_internvideo2_6b_w16_s8SlimPajama-6B-processed-8192slimpajama-6B_qwen3-8B_embedreglu2_fr_50b_jetons_fw2_6B_culturax_2B_wiki_1B_pd_1Bdetails_01-ai__Yi-6B
Dataset Card for Evaluation run of 01-ai/Yi-6B
Dataset Summary
Dataset automatically created during the evaluation run of model 01-ai/Yi-6B on the Open LLM Leaderboard.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_01-ai__Yi-6B.glove.6B.100d.txtglove.6B.100d.txt for practice
fineweb-CC-MAIN-2024-10-6B-endetails_anhnv125__pygmalion-6b-roleplay
Dataset Card for Evaluation run of anhnv125/pygmalion-6b-roleplay
Dataset Summary
Dataset automatically created during the evaluation run of model anhnv125/pygmalion-6b-roleplay on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_anhnv125__pygmalion-6b-roleplay.SlimPajama-6B-embedded
Dataset Card for SlimPajama-6B-embedded
This is a copy of DKYoon/SlimPajama-6B, together with embeddings generated by thenlper/gte-large.
There are 5.49 million examples of text, a representative random sample of SlimPajama-627B. Each text is associated with a 1024-dimensional embedding vector that is meant to represent the semantic content. The vectors were generated by average-pooling (max-pooling dataset to come in the future).
This dataset is intended to help with downstream… See the full description on the dataset page: https://huggingface.co/datasets/sproos/SlimPajama-6B-embedded.c4_pt_arrow_6bglm53-flash-fidelity-exl3-tr3-6bpw-v1
fidelity--glm53flash.malaiwah.quant.tr3-6bpw
A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/GLM-5.3-Flash-TR3-6bpw.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm53-flash-fidelity-exl3-tr3-6bpw-v1.details_KoboldAI__OPT-6B-nerys-v2
Dataset Card for Evaluation run of KoboldAI/OPT-6B-nerys-v2
Dataset Summary
Dataset automatically created during the evaluation run of model KoboldAI/OPT-6B-nerys-v2 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_KoboldAI__OPT-6B-nerys-v2.details_PygmalionAI__pygmalion-6b
Dataset Card for Evaluation run of PygmalionAI/pygmalion-6b
Dataset Summary
Dataset automatically created during the evaluation run of model PygmalionAI/pygmalion-6b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_PygmalionAI__pygmalion-6b.glove.6B.50dSlimPajama-6B_km_4_8_cos-d512details_stabilityai__stablelm-2-1_6b
Dataset Card for Evaluation run of stabilityai/stablelm-2-1_6b
Dataset automatically created during the evaluation run of model stabilityai/stablelm-2-1_6b on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_stabilityai__stablelm-2-1_6b.details_Salesforce__codegen-6B-nl
Dataset Card for Evaluation run of Salesforce/codegen-6B-nl
Dataset Summary
Dataset automatically created during the evaluation run of model Salesforce/codegen-6B-nl on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Salesforce__codegen-6B-nl.6b46ae87details_itsliupeng__fly_6btransmla_pretrain_6B_tokensdetails_bertin-project__bertin-gpt-j-6B-alpaca
Dataset Card for Evaluation run of bertin-project/bertin-gpt-j-6B-alpaca
Dataset Summary
Dataset automatically created during the evaluation run of model bertin-project/bertin-gpt-j-6B-alpaca on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bertin-project__bertin-gpt-j-6B-alpaca.SlimPajama-6B-chunked-32768-tokenized-llama3details_itsliupeng__test_6bdetails_prince-canuma__Llama-3-6B-v0.1Multilingual-BioASQ-6B
Mutilingual BioASQ-6B
We translate the BioASQ-6B English Question Answering dataset to generate parallel French, Italian and Spanish versions using the NLLB200 3B parameter model. For more info read the original task description: [http://bioasq.org/participate/challenges_year_6](http://bioasq.org/participate/challenges_year_6)
We translate the body, snippets, ideal_answer and exact_answer fields. We have validated the quality of the ideal_answer field, however, the… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Multilingual-BioASQ-6B.
