CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceFW /fineweb-edu-score-2 📚 FineWeb-Edu-score-2 1.3 trillion tokens of the finest educational data the 🌐 web has to offer What is it? 📚 FineWeb-Edu dataset consists of 1.3T tokens (FineWeb-Edu) and 5.4T tokens of educational web pages filtered from 🍷 FineWeb dataset. This is the 5.4 trillion version. Note: this version uses a lower educational score threshold = 2, which results in more documents, but lower quality compared to the 1.3T version. For more details check the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu-score-2.tabulartext-generation10B<n<100B89 likes21k downloads1y agoHugging Face02songlab /gpn-star-scores GPN-Star genome-wide scores Genome-wide, mutation-rate-calibrated GPN-Star constraint and variant scores for eight score sets covering human, mouse, chicken, D. melanogaster, C. elegans, and A. thaliana. Canonical scores are available as chromosome-sharded Parquet. Hugging Face hosts 72 BigWigs; a multi-assembly UCSC track hub references the 64 logo/LLR views. Overview and quick links Resource Link Files Browse all dataset files Default Dataset… See the full description on the dataset page: https://huggingface.co/datasets/songlab/gpn-star-scores.tabular10B<n<100B4 likes17k downloads2mo agoHugging Face03agents-course /unit4-students-scorestext10K<n<100K20 likes14k downloads1h agoHugging Face04OpenDataArena /OpenDataArena-scored-data-2603 OpenDataArena-scored-data-2603 This repository provides a scored SFT dataset collection currently featuring 63 high-quality instruction-following datasets with nearly 25 million samples. The core value lies in its 30-dimensional scoring: every sample has been evaluated on metrics such as IFD, PPL, Deita_Quality, and 27 others, enabling fine-grained data selection for filtering, curriculum learning, and mixture optimization. Key features: 30 metrics per sample — From lexical… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/OpenDataArena-scored-data-2603.text10M<n<100M9 likes6.2k downloads4mo agoHugging Face05kenhktsui /TM-DATA_quality_score_v1 Dataset Card for "TM-DATA_quality_score_v1" Adding quality score v1 to Locutusque/TM-DATA More Information needed text1M<n<10M0 likes2.5k downloads3y agoHugging Face06everycure /matrix-scores Dataset Description This dataset contains the output of the MATRIX pipeline — Every Cure's computational drug repurposing scoring system. It provides ML-generated treatment probability scores for ~39.5 million drug-disease pairs, covering ~1,800 drugs × ~22,000 diseases. ⚠️ Research use only. These scores are the output of a computational research pipeline and do not constitute medical advice, clinical recommendations, or endorsement of any drug for any use. All findings… See the full description on the dataset page: https://huggingface.co/datasets/everycure/matrix-scores.tabular10M<n<100M2 likes2.4k downloads3mo agoHugging Face07kenhktsui /cosmopedia_quality_score_v2Adding quality score v2 to HuggingFaceTB/cosmopedia tabular10M<n<100M0 likes1.9k downloads2y agoHugging Face08IlyaGusev /saiga_scoredSFT dataset for the Saiga family of models collected from various sources. tabular10K<n<100K23 likes1.6k downloads2y agoHugging Face09JanSchTech /starcoderdata-python-edu-lang-score Dataset Card for Starcoder Data with Python Education and Language Scores Dataset Summary The starcoderdata-python-edu-lang-score dataset contains the Python subset of the starcoderdata dataset. It augments the existing Python subset with features that assess the educational quality of code and classify the language of code comments. This dataset was created for high-quality Python education and language-based training, with a primary focus on facilitating models that can… See the full description on the dataset page: https://huggingface.co/datasets/JanSchTech/starcoderdata-python-edu-lang-score.tabular1M<n<10M2 likes1.5k downloads2y agoHugging Face10kenhktsui /openwebtext_quality_score_v1 Dataset Card for "openwebtext_quality_score_v1" Adding quality score v1 to Skylion007/openwebtext More Information needed texttext-generation1M<n<10M0 likes1.4k downloads3y agoHugging Face11hpprc /reranker-scores Reranker-Scores 既存の日本語検索・QAデータセットについて、データセット中のクエリに付与された正・負例の関連度を多言語・日本語reranker 5種類を用いてスコア付けしたデータセットです。クエリごとに200件程度の正・負例文書が付与されています (事例ごとに付与個数はバラバラなのでご注意ください)。score.pos.avg および score.neg.avg はそれぞれ正例・負例についての5種のrerankerの平均スコアになっています。 Short Name Hub ID bge BAAI/bge-reranker-v2-m3 gte Alibaba-NLP/gte-multilingual-reranker-base ruri cl-nagoya/ruri-reranker-large ruriv3-preview cl-nagoya/ruri-v3-reranker-310m-preview ja-ce… See the full description on the dataset page: https://huggingface.co/datasets/hpprc/reranker-scores.text100K<n<1M4 likes1.3k downloads1y agoHugging Face12zzsi /deep-scores-v2 DeepScoresV2 — Complete A HuggingFace-formatted mirror of the complete version of the DeepScoresV2 dataset for music object detection. Dataset description DeepScoresV2 is a large-scale dataset of synthetically rendered music score pages annotated with bounding boxes for musical symbols. The complete version contains 255,385 images with 151 million annotated instances across 135 symbol classes. Each image is a full score page rendered from MuseScore across 5 music fonts… See the full description on the dataset page: https://huggingface.co/datasets/zzsi/deep-scores-v2.imageobject-detection100K<n<1M0 likes867 downloads7mo agoHugging Face13iDRAMALab /iDRAMA-scored-2024 Dataset Summary iDRAMA-Scored-2024 is a large-scale dataset containing approximately 57 million social media posts from web communities on social media platform, Scored. Scored serves as an alternative to Reddit, hosting banned fringe communities, for example, c/TheDonald, a prominent right-wing community, and c/GreatAwakening, a conspiratorial community. This dataset contains 57M posts from over 950 communities collected over four years, and includes sentence embeddings for all… See the full description on the dataset page: https://huggingface.co/datasets/iDRAMALab/iDRAMA-scored-2024.tabular10M<n<100M1 likes829 downloads2y agoHugging Face14ljnlonoljpiljm /stockimage-scored-pt12image1M<n<10M0 likes713 downloads1y agoHugging Face15kenhktsui /refinedweb-3m_quality_score_v1 Dataset Card for "refinedweb-3m_quality_score_v1" Adding quality score v1 to mattymchen/refinedweb-3m More Information needed texttext-generation1M<n<10M0 likes666 downloads3y agoHugging Face16hotchpotch /hpprc_emb_reranker_score ⚠️ お知らせ よりスコア付したデータ件数とrerankerのバリエーションを増やしたデータセットのhotchpotch/hpprc_emb-scoresも公開しています。 hpprc/emb (便利なデータセットの公開、ありがとうございます)の collection と dataset がペアになっているデータに対し、negative を最大32個ランダムサンプリングしたものを、hotchpotch/japanese-bge-reranker-v2-m3-v1でスコア付けしたものです。 ライセンスは、subset ごとに hpprc/emb に記載のライセンスと同等とします。 スコア作成タイミングの revision に対してスコアを付与しているため、revision を変えると場合によって行ズレやデータ構造の変化が発生する可能性があることに注意が必要です。 例 from datasets import load_dataset # targets = ("auto-wiki-qa", "4feb2e2492")… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/hpprc_emb_reranker_score.tabular1M<n<10M4 likes643 downloads2y agoHugging Face17OpenDataArena /OpenDataArena-scored-data OpenDataArena-scored-data This repository contains a special scored and enhanced collection of over 47 original datasets for Supervised Fine-Tuning (SFT). These datasets were processed using the OpenDataArena-Tool, a comprehensive suite of automated evaluation methods for assessing instruction-following datasets. The scored-dataset collection from the OpenDataArena (ODA) reflects a large-scale, cross-domain effort to quantify dataset value in the era of large language models. ODA… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/OpenDataArena-scored-data.text10M<n<100M12 likes632 downloads5mo agoHugging Face18maveriq /fineweb-edu-bottom1pct-scoretext10M<n<100M0 likes588 downloads2y agoHugging Face19ljnlonoljpiljm /stockimage-scored-pt6image1M<n<10M0 likes578 downloads1y agoHugging Face20ljnlonoljpiljm /stockimage-scored-pt8image1M<n<10M0 likes559 downloads1y agoHugging Face21maveriq /fineweb-edu-bottom1pct-lang-scoretext10M<n<100M0 likes553 downloads2y agoHugging Face22ljnlonoljpiljm /stockimage-scored-pt1image1M<n<10M0 likes553 downloads1y agoHugging Face23ljnlonoljpiljm /BIGstockimage-1.5M-scored-pt-twoimage100K<n<1M0 likes541 downloads1y agoHugging Face24hotchpotch /mmarco-hard-negatives-reranker-score hotchpotch/mmarco-hard-negatives-reranker-score This repository contains data from mMARCO scored using the reranker BAAI/bge-reranker-v2-m3. Languages Covered target_languages = [ "english", "chinese", "french", "german", "indonesian", "italian", "portuguese", "russian", "spanish", "arabic", "dutch", "hindi", "japanese", "vietnamese" ] Hard Negative Data The hard negative data is derived from… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/mmarco-hard-negatives-reranker-score.1M<n<10M1 likes532 downloads2y agoHugging Face25ljnlonoljpiljm /stockimage-scored-pt9image1M<n<10M0 likes510 downloads1y agoHugging Face26maveriq /fineweb-edu-top1pct-lang-scoretext10M<n<100M0 likes507 downloads2y agoHugging Face27maveriq /fineweb-edu-top1pct-scoretext10M<n<100M0 likes507 downloads2y agoHugging Face28ljnlonoljpiljm /stockimage-scored-pt10image1M<n<10M0 likes495 downloads1y agoHugging Face29ljnlonoljpiljm /BIGstockimage-1.5M-scored-pt-oneimage100K<n<1M1 likes492 downloads1y agoHugging Face30KhaledReda /pairs_with_scores_v27text100M<n<1B0 likes489 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.