CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DKYoon /SlimPajama-6BSampled version of cerebras/SlimPajama-627B. Since the original data was shuffled before chunking, I only downloaded train/chunk1 (of 10 total) and further sampled 10%. This should result in roughly 6B tokens, hence SlimPajama-6B. The dataset is 24GBs in storage size when decompressed (original dataset is over 2TBs) and has 5489000 rows. The validation set and test set were sampled as well. Data source proportions for SlimPajama-627B and SlimPajama-6B For sanity purpose, I… See the full description on the dataset page: https://huggingface.co/datasets/DKYoon/SlimPajama-6B.texttext-generation1M<n<10M66 likes12k downloads3y agoHugging Face02cg1177 /hacs_segment_internvideo2_6b_w16_s80 likes1.8k downloads3y agoHugging Face03open-llm-leaderboard-old /details_EleutherAI__gpt-j-6b Dataset Card for Evaluation run of EleutherAI/gpt-j-6b Dataset Summary Dataset automatically created during the evaluation run of model EleutherAI/gpt-j-6b on the Open LLM Leaderboard. The dataset is composed of 122 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 8 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_EleutherAI__gpt-j-6b.0 likes1.6k downloads3y agoHugging Face04Yuehao /SlimPajama-6B_km-ip-d512tabular1M<n<10M0 likes813 downloads9mo agoHugging Face05minhnguyent546 /ClimbMix-6BT ClimbMix-6BT This is the tokenized nvidia/Nemotron-ClimbMix (10M subset) using SmolLM2-135M tokenzier. Data is divided into shards (.npy files) for easier to load with PyTorch IterableDataset. Each .npy file can be loaded with numpy.load('file_name.npy'). Split # Documents # Shards # Tokens train 9,900,000 65 6,463,974,020 (6.5B) val 100,000 1 64,859,672 (65M) Total 10,000,000 66 6,528,833,692 (6.5B) Example of usage uvx hf download… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/ClimbMix-6BT.text-generation0 likes714 downloads6mo agoHugging Face06cg1177 /activitynet_internvideo2_6b_w16_s83 likes636 downloads3y agoHugging Face07tj-solergibert /SlimPajama-6B-processed-8192100K<n<1M0 likes635 downloads3y agoHugging Face08Yuehao /slimpajama-6B_qwen3-8B_embed1M<n<10M0 likes617 downloads10mo agoHugging Face09saifgazali /reglu2_fr_50b_jetons_fw2_6B_culturax_2B_wiki_1B_pd_1Btext10M<n<100M0 likes526 downloads4mo agoHugging Face10open-llm-leaderboard-old /details_01-ai__Yi-6B Dataset Card for Evaluation run of 01-ai/Yi-6B Dataset Summary Dataset automatically created during the evaluation run of model 01-ai/Yi-6B on the Open LLM Leaderboard. The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_01-ai__Yi-6B.1 likes392 downloads3y agoHugging Face11SLU-CSCI4750 /glove.6B.100d.txtglove.6B.100d.txt for practice text100K<n<1M0 likes376 downloads3y agoHugging Face12prince-canuma /fineweb-CC-MAIN-2024-10-6B-entabular1M<n<10M0 likes312 downloads2y agoHugging Face13open-llm-leaderboard-old /details_anhnv125__pygmalion-6b-roleplay Dataset Card for Evaluation run of anhnv125/pygmalion-6b-roleplay Dataset Summary Dataset automatically created during the evaluation run of model anhnv125/pygmalion-6b-roleplay on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_anhnv125__pygmalion-6b-roleplay.0 likes288 downloads3y agoHugging Face14sproos /SlimPajama-6B-embedded Dataset Card for SlimPajama-6B-embedded This is a copy of DKYoon/SlimPajama-6B, together with embeddings generated by thenlper/gte-large. There are 5.49 million examples of text, a representative random sample of SlimPajama-627B. Each text is associated with a 1024-dimensional embedding vector that is meant to represent the semantic content. The vectors were generated by average-pooling (max-pooling dataset to come in the future). This dataset is intended to help with downstream… See the full description on the dataset page: https://huggingface.co/datasets/sproos/SlimPajama-6B-embedded.text1M<n<10M3 likes276 downloads3y agoHugging Face15jungsin3 /c4_pt_arrow_6b1M<n<10M0 likes237 downloads7mo agoHugging Face16malaiwah /glm53-flash-fidelity-exl3-tr3-6bpw-v1 fidelity--glm53flash.malaiwah.quant.tr3-6bpw A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/GLM-5.3-Flash-TR3-6bpw. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm53-flash-fidelity-exl3-tr3-6bpw-v1.tabularn<1K0 likes209 downloads20d agoHugging Face17open-llm-leaderboard-old /details_KoboldAI__OPT-6B-nerys-v2 Dataset Card for Evaluation run of KoboldAI/OPT-6B-nerys-v2 Dataset Summary Dataset automatically created during the evaluation run of model KoboldAI/OPT-6B-nerys-v2 on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_KoboldAI__OPT-6B-nerys-v2.0 likes204 downloads3y agoHugging Face18open-llm-leaderboard-old /details_PygmalionAI__pygmalion-6b Dataset Card for Evaluation run of PygmalionAI/pygmalion-6b Dataset Summary Dataset automatically created during the evaluation run of model PygmalionAI/pygmalion-6b on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_PygmalionAI__pygmalion-6b.1 likes198 downloads3y agoHugging Face19antokun /glove.6B.50dtext100K<n<1M0 likes187 downloads2y agoHugging Face20Yuehao /SlimPajama-6B_km_4_8_cos-d512tabular1M<n<10M0 likes185 downloads8mo agoHugging Face21open-llm-leaderboard-old /details_stabilityai__stablelm-2-1_6b Dataset Card for Evaluation run of stabilityai/stablelm-2-1_6b Dataset automatically created during the evaluation run of model stabilityai/stablelm-2-1_6b on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_stabilityai__stablelm-2-1_6b.0 likes181 downloads3y agoHugging Face22open-llm-leaderboard-old /details_Salesforce__codegen-6B-nl Dataset Card for Evaluation run of Salesforce/codegen-6B-nl Dataset Summary Dataset automatically created during the evaluation run of model Salesforce/codegen-6B-nl on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Salesforce__codegen-6B-nl.0 likes171 downloads3y agoHugging Face23ectihex /6b46ae870 likes160 downloads1mo agoHugging Face24open-llm-leaderboard-old /details_itsliupeng__fly_6b0 likes148 downloads2y agoHugging Face25fxmeng /transmla_pretrain_6B_tokenstext1M<n<10M0 likes147 downloads1y agoHugging Face26open-llm-leaderboard-old /details_bertin-project__bertin-gpt-j-6B-alpaca Dataset Card for Evaluation run of bertin-project/bertin-gpt-j-6B-alpaca Dataset Summary Dataset automatically created during the evaluation run of model bertin-project/bertin-gpt-j-6B-alpaca on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bertin-project__bertin-gpt-j-6B-alpaca.0 likes146 downloads3y agoHugging Face27sheryc /SlimPajama-6B-chunked-32768-tokenized-llama310K<n<100K0 likes146 downloads2y agoHugging Face28open-llm-leaderboard-old /details_itsliupeng__test_6b0 likes137 downloads2y agoHugging Face29open-llm-leaderboard-old /details_prince-canuma__Llama-3-6B-v0.10 likes137 downloads2y agoHugging Face30HiTZ /Multilingual-BioASQ-6B Mutilingual BioASQ-6B We translate the BioASQ-6B English Question Answering dataset to generate parallel French, Italian and Spanish versions using the NLLB200 3B parameter model. For more info read the original task description: [http://bioasq.org/participate/challenges_year_6](http://bioasq.org/participate/challenges_year_6) We translate the body, snippets, ideal_answer and exact_answer fields. We have validated the quality of the ideal_answer field, however, the… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Multilingual-BioASQ-6B.textquestion-answering10K<n<100K2 likes130 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.