CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01songlab /gpn-star-scores GPN-Star genome-wide scores Genome-wide, mutation-rate-calibrated GPN-Star constraint and variant scores for eight score sets covering human, mouse, chicken, D. melanogaster, C. elegans, and A. thaliana. Canonical scores are available as chromosome-sharded Parquet. Hugging Face hosts 72 BigWigs; a multi-assembly UCSC track hub references the 64 logo/LLR views. Overview and quick links Resource Link Files Browse all dataset files Default Dataset… See the full description on the dataset page: https://huggingface.co/datasets/songlab/gpn-star-scores.tabular10B<n<100B4 likes17k downloads2mo agoHugging Face02agents-course /unit4-students-scorestext10K<n<100K20 likes14k downloads3h agoHugging Face03JQL-AI /hplt2_edu_scores HPLT2-Edu-scores Dataset summary HPLT2-JQL-Education is a model-annotated language subset of HPLT2, spanning 35 languages. Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction. The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance. For example, in the Spanish language case… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/hplt2_edu_scores.tabulartext-ranking1B<n<10B1 likes6k downloads1y agoHugging Face04JQL-AI /fw2_edu_scores Fineweb2-Edu-scores Dataset summary FineWeb2-JQL-Education is a model-annotated language subset of FineWeb2, spanning 36 languages. Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction. The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance. For example, in the Spanish language… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/fw2_edu_scores.tabulartext-ranking1B<n<10B6 likes4k downloads1y agoHugging Face05everycure /matrix-scores Dataset Description This dataset contains the output of the MATRIX pipeline — Every Cure's computational drug repurposing scoring system. It provides ML-generated treatment probability scores for ~39.5 million drug-disease pairs, covering ~1,800 drugs × ~22,000 diseases. ⚠️ Research use only. These scores are the output of a computational research pipeline and do not constitute medical advice, clinical recommendations, or endorsement of any drug for any use. All findings… See the full description on the dataset page: https://huggingface.co/datasets/everycure/matrix-scores.tabular10M<n<100M2 likes2.4k downloads3mo agoHugging Face06alvations /encoder-decoder-floresp-scores0 likes2.4k downloads2y agoHugging Face07GEM-submissions /v2-outputs-and-scores0 likes2.2k downloads3y agoHugging Face08ngeorgea /transfermarkt-player-scores Transfermarkt Player Scores (mirror) Зеркало transfermarkt-datasets — структурированные футбольные данные из Transfermarkt. Источник: davidcariboo/player-scores / R2 CDN Таблицы (12) competitions, clubs, players, games, appearances, player_valuations, club_games, game_events, game_lineups, transfers, countries, national_teams Использование from datasets import load_dataset players = load_dataset("ngeorgea/transfermarkt-player-scores"… See the full description on the dataset page: https://huggingface.co/datasets/ngeorgea/transfermarkt-player-scores.tabular-classification0 likes1.4k downloads2mo agoHugging Face09hpprc /reranker-scores Reranker-Scores 既存の日本語検索・QAデータセットについて、データセット中のクエリに付与された正・負例の関連度を多言語・日本語reranker 5種類を用いてスコア付けしたデータセットです。クエリごとに200件程度の正・負例文書が付与されています (事例ごとに付与個数はバラバラなのでご注意ください)。score.pos.avg および score.neg.avg はそれぞれ正例・負例についての5種のrerankerの平均スコアになっています。 Short Name Hub ID bge BAAI/bge-reranker-v2-m3 gte Alibaba-NLP/gte-multilingual-reranker-base ruri cl-nagoya/ruri-reranker-large ruriv3-preview cl-nagoya/ruri-v3-reranker-310m-preview ja-ce… See the full description on the dataset page: https://huggingface.co/datasets/hpprc/reranker-scores.text100K<n<1M4 likes1.3k downloads1y agoHugging Face10ShantanuD /Temporal_Awareness_Node_Scores0 likes1.3k downloads5mo agoHugging Face11jacobmorrison /scores_13099 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_13099', 'hf_repo_id_scores': 'scores_13099', 'input_filename': '/output/shards/13099/237.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths':… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/scores_13099.text1M<n<10M0 likes1.2k downloads2y agoHugging Face12kaushikreddyxyz /corpus-scores-dom-layer8 corpus-scores-dom-layer8 Difference-of-means (DoM) STEERING scores for ClimbMix, gemma-2-2b layer 8. Kept separate from the detection score store corpus-scores (overflow). Files (per shard <sid>, ClimbMix shards 320-362, 43 total) scores_<sid>.npy — int8 [n_tokens, 54]; axis 0 token, axis 1 concept (index into columns.json concepts[]) tokens_<sid>.npy — int32 [n_tokens] gemma tokenizer ids docs_<sid>.jsonl — per-document metadata (token spans)… See the full description on the dataset page: https://huggingface.co/datasets/kaushikreddyxyz/corpus-scores-dom-layer8.1B<n<10B0 likes1.1k downloads3mo agoHugging Face13songlab /gpn-msa-hg38-scores GPN-MSA predictions for all possible SNPs in the human genome (~9 billion) For more information check out our paper and repository. Querying specific variants or genes Install the latest tabix:In your current conda environment (might be slow):conda install -c bioconda -c conda-forge htslib=1.18 or in a new conda environment:conda create -n tabix -c bioconda -c conda-forge htslib=1.18 conda activate tabix Query a specific region (e.g. BRCA1), from the remote file:… See the full description on the dataset page: https://huggingface.co/datasets/songlab/gpn-msa-hg38-scores.5 likes933 downloads2y agoHugging Face14jacobmorrison /scores_4036 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_4036', 'hf_repo_id_scores': 'scores_4036', 'input_filename': '/output/shards/4036/27.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths': ['/reward_model'], 'num_completions':… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/scores_4036.text1M<n<10M1 likes931 downloads2y agoHugging Face15zzsi /deep-scores-v2 DeepScoresV2 — Complete A HuggingFace-formatted mirror of the complete version of the DeepScoresV2 dataset for music object detection. Dataset description DeepScoresV2 is a large-scale dataset of synthetically rendered music score pages annotated with bounding boxes for musical symbols. The complete version contains 255,385 images with 151 million annotated instances across 135 symbol classes. Each image is a full score page rendered from MuseScore across 5 music fonts… See the full description on the dataset page: https://huggingface.co/datasets/zzsi/deep-scores-v2.imageobject-detection100K<n<1M0 likes867 downloads7mo agoHugging Face16eco4cast /neon4cast-scoresSnapshot of the Ecological Forecasting Initiative NEON Forecasting Challenge Includes probabilistic forecasts, observations, and skill scores across all submitted forecasts over 5 challenge themes. tabularn<1K0 likes834 downloads3y agoHugging Face17manu /reranker-scorestabularn<1K0 likes738 downloads3y agoHugging Face18kaushikreddyxyz /corpus-scores corpus-scores Concept-probe detection scores for ClimbMix, layers 6/8/14 co-located per token. Shards here: 320-355. Overflow (shards 356-362): kaushikreddyxyz/corpus-scores-overflow. Files (per shard <sid>, ClimbMix shards 320-362) scores_<sid>.npy — int8 [n_tokens, 3, 54] axis 0: token (aligns 1:1 with tokens_<sid>.npy / docs_<sid>.jsonl) axis 1: layer — 0 = layer 6, 1 = layer 8, 2 = layer 14 axis 2: concept — index into columns.json concepts[] (54 concepts, 7… See the full description on the dataset page: https://huggingface.co/datasets/kaushikreddyxyz/corpus-scores.1B<n<10B0 likes722 downloads2mo agoHugging Face19Eurolingua /hplt3_edu_scores HPLT3-Edu-scores Dataset summary HPLT3-JQL-Education is a model-annotated language subset of HPLT3, spanning 36 languages. Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction. HPLT3-Edu-scores was created based on scores assigned by a deep learning classifier trained to identify educational samples using Snowflake's Arctic-embed-m-v2.0 embeddings. For all training ablations, we used… See the full description on the dataset page: https://huggingface.co/datasets/Eurolingua/hplt3_edu_scores.tabulartext-ranking1B<n<10B0 likes596 downloads6mo agoHugging Face20jacobmorrison /scores_6511 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_6511', 'hf_repo_id_scores': 'scores_6511', 'input_filename': '/output/shards/6511/24.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths': ['Skywork/Skywork-Reward-Llama-3.1-8B']… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/scores_6511.text100K<n<1M0 likes538 downloads2y agoHugging Face21JQL-AI /curated_edu_scorestabularn<1K0 likes505 downloads1y agoHugging Face22osyvokon /pavlick-formality-scoresThis dataset contains sentence-level formality annotations used in the 2016 TACL paper "An Empirical Analysis of Formality in Online Communication" (Pavlick and Tetreault, 2016). It includes sentences from four genres (news, blogs, email, and QA forums), all annotated by humans on Amazon Mechanical Turk. The news and blog data was collected by Shibamouli Lahiri, and we are redistributing it here for the convenience of other researchers. We collected the email and answers data ourselves, using… See the full description on the dataset page: https://huggingface.co/datasets/osyvokon/pavlick-formality-scores.texttext-classification10K<n<100K4 likes503 downloads3y agoHugging Face23KhaledReda /pairs_with_scores_v27text100M<n<1B0 likes489 downloads8mo agoHugging Face24Finnish-NLP /swe_Finepdfs_edu_scorestext1M<n<10M0 likes464 downloads11mo agoHugging Face25SultanR /AraMix-Translation-Scores AraMix-Translation-Scores AdaMLLab/AraMix (minhash_deduped subset, 178,883,241 rows) with a machine-translation-detection score added to every document. All original columns are preserved. Columns column type description id string unchanged from AraMix source string unchanged from AraMix text string unchanged from AraMix mmbert_quality_score float64 AraMix's original mmbert_score, renamed mmbert_translated_score float64 new —… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/AraMix-Translation-Scores.tabular100M<n<1B0 likes385 downloads2mo agoHugging Face26ntranoslab /vesm_scores Proteome-wide VESM variant effect scores This repository provides precomputed proteome-wide (UniProtKB, hg19, and hg38) variant-effect prediction scores using the latest VESM models developed in the paper "Compressing the collective knowledge of ESM into a single protein language model" by Tuan Dinh, Seon-Kyeong Jang, Noah Zaitlen and Vasilis Ntranos. Models: VESM_3B, VESM3, sequence-only VESM3, and VESM++ (available at https://huggingface.co/ntranoslab/vesm). VESM_3B and VESM3… See the full description on the dataset page: https://huggingface.co/datasets/ntranoslab/vesm_scores.tabular1B<n<10B2 likes380 downloads1y agoHugging Face27Finnish-NLP /dan_Fineweb2_edu_scorestext10M<n<100M0 likes357 downloads11mo agoHugging Face28hotchpotch /japanese-reranker-v2-hard-negatives-scores hotchpotch/japanese-reranker-v2-hard-negatives-scores This dataset was used to train the Japanese Reranker v2 family. It was not generated by Japanese Reranker v2 models. Unified hard-negative score rows generated from the training sources used for Japanese Reranker v2-family experiments. This dataset keeps teacher scores as floating-point soft labels. It follows the HPPRC Emb Score style: score/id rows and referenced document text collections are stored separately. Each score… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/japanese-reranker-v2-hard-negatives-scores.text10M<n<100M1 likes352 downloads4mo agoHugging Face29Finnish-NLP /swe_Fineweb2_edu_scorestext10M<n<100M0 likes325 downloads11mo agoHugging Face30Finnish-NLP /nob_Fineweb2_edu_scorestext10M<n<100M0 likes323 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.