CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01opendatalab /SlimPajama-Meta-rater Annotated SlimPajama Dataset Dataset Description This dataset contains the first fully annotated SlimPajama dataset with comprehensive quality metrics for data-centric large language model research. The dataset includes approximately 580 billion tokens from the training set of the original SlimPajama dataset, annotated across 25 different quality dimensions. Note: This dataset contains only the training set portion of the original SlimPajama dataset, which is why the… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater.tabulartext-generation10M<n<100M7 likes5.2k downloads1y agoHugging Face02opendatalab /SlimPajama-Meta-rater-Readability-30B Top 30B token SlimPajama Subset selected by the Readability rater This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models. Code: https://github.com/opendatalab/Meta-rater Dataset Description This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Readability dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Readability-30B.tabulartext-generation1M<n<10M1 likes358 downloads1y agoHugging Face03SkillCorner /opendata-bodypose SkillCorner Open Data — Body Pose 3D body-pose data derived from broadcast video, released alongside the SkillCorner Open Data repository as a joint initiative between SkillCorner and PySport. Initial testing release. Two matches, published so the community can work with the format and tell us what is useful before we consider a wider release. Feedback is genuinely wanted — open an issue on the opendata repo or reply in the Community tab here. What is in here… See the full description on the dataset page: https://huggingface.co/datasets/SkillCorner/opendata-bodypose.tabular100K<n<1M0 likes350 downloads15d agoHugging Face04opendatalab /SlimPajama-Meta-rater-Professionalism-30B Top 30B token SlimPajama Subset selected by the Professionalism rater This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models. Code: https://github.com/opendatalab/Meta-rater Dataset Description This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Professionalism dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness)… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Professionalism-30B.tabulartext-generation1M<n<10M0 likes227 downloads1y agoHugging Face05opendatalab /Meta-rater-PRRC-Rater-dataset PRRC Rater Training and Evaluation Dataset Dataset Description This dataset contains the full training and evaluation data for the PRRC rater models described in Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models. It is designed for training and benchmarking models that score text along four key quality dimensions: Professionalism, Readability, Reasoning, and Cleanliness. Source: Subset of SlimPajama-627B, annotated for PRRC dimensions… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/Meta-rater-PRRC-Rater-dataset.tabulartext-classification100K<n<1M1 likes201 downloads1y agoHugging Face06opendatalab /SlimPajama-Meta-rater-Reasoning-30B Top 30B token SlimPajama Subset selected by the Reasoning rater This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models. Code: https://github.com/opendatalab/Meta-rater Dataset Description This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Reasoning dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework. Each… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Reasoning-30B.tabulartext-generation1M<n<10M1 likes186 downloads1y agoHugging Face07OpenChristianDataOrg /open-christian-data Open Christian Data Open Christian Data (OCD) aims to be the single unified collection of all public domain Christian text. It exists to bring this collection together from across the internet and to structure it in a useful format for public use. Beyond the Bible, Christian writing is poorly represented as a cohesive dataset or as data structured for AI training. This Hugging Face release is the AI-focused publication of the collection: consistent, downloadable JSON for model… See the full description on the dataset page: https://huggingface.co/datasets/OpenChristianDataOrg/open-christian-data.tabulartext-generation100K<n<1M0 likes156 downloads2mo agoHugging Face08opendatalab /SlimPajama-Meta-rater-Cleanliness-30B Top 30B token SlimPajama Subset selected by the Cleanliness rater This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models. Code: https://github.com/opendatalab/Meta-rater Dataset Description This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Cleanliness dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Cleanliness-30B.tabulartext-generation1M<n<10M0 likes101 downloads1y agoHugging Face09gohumanize /gohumanize-open-humanizer-dataset GoHumanize Open Humanizer Dataset 2,957 training pairs and 300 test pairs for teaching a language model to rewrite AI-styled English prose into natural human writing. Each pair is: input: a passage rewritten by a large language model in the register typical of LLM output (formal, smooth, hedged, connective phrases, no contractions); output: the original human-written passage, from a public-domain book or, since version 2, from a US federal government publication. The human… See the full description on the dataset page: https://huggingface.co/datasets/gohumanize/gohumanize-open-humanizer-dataset.tabulartext-generation1K<n<10K0 likes74 downloads13h agoHugging Face10open-llm-leaderboard /databricks__dolly-v2-7b-detailsgated Dataset Card for Evaluation run of databricks/dolly-v2-7b Dataset automatically created during the evaluation run of model databricks/dolly-v2-7b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-7b-details.tabular10K<n<100K0 likes59 downloads2y agoHugging Face11open-llm-leaderboard /godlikehhd__alpaca_data_score_max_0.1_2600-detailsgated Dataset Card for Evaluation run of godlikehhd/alpaca_data_score_max_0.1_2600 Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_score_max_0.1_2600 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_score_max_0.1_2600-details.tabular10K<n<100K0 likes47 downloads2y agoHugging Face12open-llm-leaderboard /databricks__dolly-v2-12b-detailsgated Dataset Card for Evaluation run of databricks/dolly-v2-12b Dataset automatically created during the evaluation run of model databricks/dolly-v2-12b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-12b-details.tabular10K<n<100K0 likes44 downloads2y agoHugging Face13open-llm-leaderboard /databricks__dbrx-base-detailsgated Dataset Card for Evaluation run of databricks/dbrx-base Dataset automatically created during the evaluation run of model databricks/dbrx-base The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dbrx-base-details.tabular1K<n<10K0 likes25 downloads2y agoHugging Face14open-llm-leaderboard /databricks__dolly-v1-6b-detailsgated Dataset Card for Evaluation run of databricks/dolly-v1-6b Dataset automatically created during the evaluation run of model databricks/dolly-v1-6b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v1-6b-details.tabular10K<n<100K0 likes23 downloads2y agoHugging Face15open-llm-leaderboard /databricks__dolly-v2-3b-detailsgated Dataset Card for Evaluation run of databricks/dolly-v2-3b Dataset automatically created during the evaluation run of model databricks/dolly-v2-3b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-3b-details.tabular10K<n<100K0 likes22 downloads2y agoHugging Face16open-llm-leaderboard /BEE-spoke-data__Meta-Llama-3-8Bee-detailsgated Dataset Card for Evaluation run of BEE-spoke-data/Meta-Llama-3-8Bee Dataset automatically created during the evaluation run of model BEE-spoke-data/Meta-Llama-3-8Bee The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BEE-spoke-data__Meta-Llama-3-8Bee-details.tabular10K<n<100K0 likes21 downloads2y agoHugging Face17open-llm-leaderboard /godlikehhd__alpaca_data_ifd_max_2600-detailsgated Dataset Card for Evaluation run of godlikehhd/alpaca_data_ifd_max_2600 Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_ifd_max_2600 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_ifd_max_2600-details.tabular10K<n<100K0 likes16 downloads2y agoHugging Face18MASSJ77 /open-dialogue-dataset-5k-sampleOpen Dialogue Dataset – 5k Cleaned Instruction-Response Conversations (Free Sample) 1️.Description This dataset contains 5,000 instruction-response dialogues sampled from the full 189k cleaned dialogues dataset. Each conversation has 2 turns: User: instruction or question Assistant: response or explanation It is fully cleaned, deduplicated, and structured in JSONL and CSV formats, ready for: Fine-tuning large language models (LLMs) Instruction-tuned chatbot training NLP research and analysis… See the full description on the dataset page: https://huggingface.co/datasets/MASSJ77/open-dialogue-dataset-5k-sample.tabulartext-generation1K<n<10K1 likes13 downloads9mo agoHugging Face19open-llm-leaderboard /BEE-spoke-data__smol_llama-220M-openhermes-detailsgated Dataset Card for Evaluation run of BEE-spoke-data/smol_llama-220M-openhermes Dataset automatically created during the evaluation run of model BEE-spoke-data/smol_llama-220M-openhermes The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BEE-spoke-data__smol_llama-220M-openhermes-details.tabular10K<n<100K0 likes11 downloads2y agoHugging Face20open-llm-leaderboard /BEE-spoke-data__smol_llama-220M-GQA-detailsgated Dataset Card for Evaluation run of BEE-spoke-data/smol_llama-220M-GQA Dataset automatically created during the evaluation run of model BEE-spoke-data/smol_llama-220M-GQA The dataset is composed of 43 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BEE-spoke-data__smol_llama-220M-GQA-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face21open-llm-leaderboard /BEE-spoke-data__smol_llama-101M-GQA-detailsgated Dataset Card for Evaluation run of BEE-spoke-data/smol_llama-101M-GQA Dataset automatically created during the evaluation run of model BEE-spoke-data/smol_llama-101M-GQA The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BEE-spoke-data__smol_llama-101M-GQA-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face22open-llm-leaderboard /godlikehhd__alpaca_data_sampled_ifd_5200-detailsgated Dataset Card for Evaluation run of godlikehhd/alpaca_data_sampled_ifd_5200 Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_sampled_ifd_5200 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_sampled_ifd_5200-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face23open-llm-leaderboard /godlikehhd__alpaca_data_ins_max_5200-detailsgated Dataset Card for Evaluation run of godlikehhd/alpaca_data_ins_max_5200 Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_ins_max_5200 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_ins_max_5200-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face24open-llm-leaderboard /OpenLLM-France__Lucie-7B-Instruct-human-data-detailsgated Dataset Card for Evaluation run of OpenLLM-France/Lucie-7B-Instruct-human-data Dataset automatically created during the evaluation run of model OpenLLM-France/Lucie-7B-Instruct-human-data The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/OpenLLM-France__Lucie-7B-Instruct-human-data-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face25open-llm-leaderboard /godlikehhd__alpaca_data_ifd_min_2600-detailsgated Dataset Card for Evaluation run of godlikehhd/alpaca_data_ifd_min_2600 Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_ifd_min_2600 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_ifd_min_2600-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face26zakerytclarke /open_food_dataset_cleanedtabular100K<n<1M0 likes9 downloads1y agoHugging Face27open-llm-leaderboard /BEE-spoke-data__smol_llama-220M-GQA-fineweb_edu-detailsgated Dataset Card for Evaluation run of BEE-spoke-data/smol_llama-220M-GQA-fineweb_edu Dataset automatically created during the evaluation run of model BEE-spoke-data/smol_llama-220M-GQA-fineweb_edu The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BEE-spoke-data__smol_llama-220M-GQA-fineweb_edu-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face28open-llm-leaderboard /godlikehhd__alpaca_data_full_2-detailsgated Dataset Card for Evaluation run of godlikehhd/alpaca_data_full_2 Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_full_2 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_full_2-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face29open-llm-leaderboard /godlikehhd__alpaca_data_ins_min_2600-detailsgated Dataset Card for Evaluation run of godlikehhd/alpaca_data_ins_min_2600 Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_ins_min_2600 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_ins_min_2600-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face30open-llm-leaderboard /godlikehhd__alpaca_data_score_max_2600_3B-detailsgated Dataset Card for Evaluation run of godlikehhd/alpaca_data_score_max_2600_3B Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_score_max_2600_3B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_score_max_2600_3B-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.