datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpn-star-scores
GPN-Star genome-wide scores
Genome-wide, mutation-rate-calibrated GPN-Star constraint and variant scores for
eight score sets covering human, mouse, chicken, D. melanogaster,
C. elegans, and A. thaliana. Canonical scores are available as
chromosome-sharded Parquet. Hugging Face hosts 72 BigWigs; a multi-assembly
UCSC track hub references the 64 logo/LLR views.
Overview and quick links
Resource
Link
Files
Browse all dataset files
Default Dataset… See the full description on the dataset page: https://huggingface.co/datasets/songlab/gpn-star-scores.unit4-students-scoreshplt2_edu_scores
HPLT2-Edu-scores
Dataset summary
HPLT2-JQL-Education is a model-annotated language subset of HPLT2, spanning 35 languages.
Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction.
The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance.
For example, in the Spanish language case… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/hplt2_edu_scores.fw2_edu_scores
Fineweb2-Edu-scores
Dataset summary
FineWeb2-JQL-Education is a model-annotated language subset of FineWeb2, spanning 36 languages.
Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction.
The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance.
For example, in the Spanish language… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/fw2_edu_scores.matrix-scores
Dataset Description
This dataset contains the output of the MATRIX pipeline — Every Cure's computational drug repurposing scoring system. It provides ML-generated treatment probability scores for ~39.5 million drug-disease pairs, covering ~1,800 drugs × ~22,000 diseases.
⚠️ Research use only. These scores are the output of a computational research pipeline and do not constitute medical advice, clinical recommendations, or endorsement of any drug for any use. All findings… See the full description on the dataset page: https://huggingface.co/datasets/everycure/matrix-scores.encoder-decoder-floresp-scoresv2-outputs-and-scorestransfermarkt-player-scores
Transfermarkt Player Scores (mirror)
Зеркало transfermarkt-datasets — структурированные футбольные данные из Transfermarkt.
Источник: davidcariboo/player-scores / R2 CDN
Таблицы (12)
competitions, clubs, players, games, appearances, player_valuations, club_games, game_events, game_lineups, transfers, countries, national_teams
Использование
from datasets import load_dataset
players = load_dataset("ngeorgea/transfermarkt-player-scores"… See the full description on the dataset page: https://huggingface.co/datasets/ngeorgea/transfermarkt-player-scores.reranker-scores
Reranker-Scores
既存の日本語検索・QAデータセットについて、データセット中のクエリに付与された正・負例の関連度を多言語・日本語reranker 5種類を用いてスコア付けしたデータセットです。クエリごとに200件程度の正・負例文書が付与されています (事例ごとに付与個数はバラバラなのでご注意ください)。score.pos.avg および score.neg.avg はそれぞれ正例・負例についての5種のrerankerの平均スコアになっています。
Short Name
Hub ID
bge
BAAI/bge-reranker-v2-m3
gte
Alibaba-NLP/gte-multilingual-reranker-base
ruri
cl-nagoya/ruri-reranker-large
ruriv3-preview
cl-nagoya/ruri-v3-reranker-310m-preview
ja-ce… See the full description on the dataset page: https://huggingface.co/datasets/hpprc/reranker-scores.Temporal_Awareness_Node_Scoresscores_13099
allenai/open_instruct: Rejection Sampling Dataset
See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail
Configs
args:
{'add_timestamp': False,
'hf_entity': 'jacobmorrison',
'hf_repo_id': 'rejection_sampling_13099',
'hf_repo_id_scores': 'scores_13099',
'input_filename': '/output/shards/13099/237.jsonl',
'max_forward_batch_size': 64,
'mode': 'judgement',
'model_names_or_paths':… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/scores_13099.corpus-scores-dom-layer8
corpus-scores-dom-layer8
Difference-of-means (DoM) STEERING scores for ClimbMix, gemma-2-2b layer 8.
Kept separate from the detection score store
corpus-scores
(overflow).
Files (per shard <sid>, ClimbMix shards 320-362, 43 total)
scores_<sid>.npy — int8 [n_tokens, 54]; axis 0 token, axis 1 concept
(index into columns.json concepts[])
tokens_<sid>.npy — int32 [n_tokens] gemma tokenizer ids
docs_<sid>.jsonl — per-document metadata (token spans)… See the full description on the dataset page: https://huggingface.co/datasets/kaushikreddyxyz/corpus-scores-dom-layer8.gpn-msa-hg38-scores
GPN-MSA predictions for all possible SNPs in the human genome (~9 billion)
For more information check out our paper and repository.
Querying specific variants or genes
Install the latest tabix:In your current conda environment (might be slow):conda install -c bioconda -c conda-forge htslib=1.18
or in a new conda environment:conda create -n tabix -c bioconda -c conda-forge htslib=1.18
conda activate tabix
Query a specific region (e.g. BRCA1), from the remote file:… See the full description on the dataset page: https://huggingface.co/datasets/songlab/gpn-msa-hg38-scores.scores_4036
allenai/open_instruct: Rejection Sampling Dataset
See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail
Configs
args:
{'add_timestamp': False,
'hf_entity': 'jacobmorrison',
'hf_repo_id': 'rejection_sampling_4036',
'hf_repo_id_scores': 'scores_4036',
'input_filename': '/output/shards/4036/27.jsonl',
'max_forward_batch_size': 64,
'mode': 'judgement',
'model_names_or_paths': ['/reward_model'],
'num_completions':… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/scores_4036.deep-scores-v2
DeepScoresV2 — Complete
A HuggingFace-formatted mirror of the complete version of the
DeepScoresV2 dataset for music object detection.
Dataset description
DeepScoresV2 is a large-scale dataset of synthetically rendered music score pages
annotated with bounding boxes for musical symbols. The complete version contains
255,385 images with 151 million annotated instances across 135 symbol classes.
Each image is a full score page rendered from MuseScore across 5 music fonts… See the full description on the dataset page: https://huggingface.co/datasets/zzsi/deep-scores-v2.neon4cast-scoresSnapshot of the Ecological Forecasting Initiative NEON Forecasting Challenge
Includes probabilistic forecasts, observations, and skill scores across all submitted forecasts over 5 challenge themes.
reranker-scorescorpus-scores
corpus-scores
Concept-probe detection scores for ClimbMix, layers 6/8/14 co-located per token.
Shards here: 320-355. Overflow (shards 356-362): kaushikreddyxyz/corpus-scores-overflow.
Files (per shard <sid>, ClimbMix shards 320-362)
scores_<sid>.npy — int8 [n_tokens, 3, 54]
axis 0: token (aligns 1:1 with tokens_<sid>.npy / docs_<sid>.jsonl)
axis 1: layer — 0 = layer 6, 1 = layer 8, 2 = layer 14
axis 2: concept — index into columns.json concepts[] (54 concepts, 7… See the full description on the dataset page: https://huggingface.co/datasets/kaushikreddyxyz/corpus-scores.hplt3_edu_scores
HPLT3-Edu-scores
Dataset summary
HPLT3-JQL-Education is a model-annotated language subset of HPLT3, spanning 36 languages.
Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction.
HPLT3-Edu-scores was created based on scores assigned by a deep learning classifier trained to identify educational samples using Snowflake's Arctic-embed-m-v2.0 embeddings.
For all training ablations, we used… See the full description on the dataset page: https://huggingface.co/datasets/Eurolingua/hplt3_edu_scores.scores_6511
allenai/open_instruct: Rejection Sampling Dataset
See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail
Configs
args:
{'add_timestamp': False,
'hf_entity': 'jacobmorrison',
'hf_repo_id': 'rejection_sampling_6511',
'hf_repo_id_scores': 'scores_6511',
'input_filename': '/output/shards/6511/24.jsonl',
'max_forward_batch_size': 64,
'mode': 'judgement',
'model_names_or_paths': ['Skywork/Skywork-Reward-Llama-3.1-8B']… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/scores_6511.curated_edu_scorespavlick-formality-scoresThis dataset contains sentence-level formality annotations used in the 2016
TACL paper "An Empirical Analysis of Formality in Online Communication"
(Pavlick and Tetreault, 2016). It includes sentences from four genres (news,
blogs, email, and QA forums), all annotated by humans on Amazon Mechanical
Turk. The news and blog data was collected by Shibamouli Lahiri, and we are
redistributing it here for the convenience of other researchers. We collected
the email and answers data ourselves, using… See the full description on the dataset page: https://huggingface.co/datasets/osyvokon/pavlick-formality-scores.pairs_with_scores_v27swe_Finepdfs_edu_scoresAraMix-Translation-Scores
AraMix-Translation-Scores
AdaMLLab/AraMix (minhash_deduped
subset, 178,883,241 rows) with a machine-translation-detection score added to every
document. All original columns are preserved.
Columns
column
type
description
id
string
unchanged from AraMix
source
string
unchanged from AraMix
text
string
unchanged from AraMix
mmbert_quality_score
float64
AraMix's original mmbert_score, renamed
mmbert_translated_score
float64
new —… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/AraMix-Translation-Scores.vesm_scores
Proteome-wide VESM variant effect scores
This repository provides precomputed proteome-wide (UniProtKB, hg19, and hg38) variant-effect prediction scores using the latest VESM models developed in the paper "Compressing the collective knowledge of ESM into a single protein language model" by Tuan Dinh, Seon-Kyeong Jang, Noah Zaitlen and Vasilis Ntranos.
Models: VESM_3B, VESM3, sequence-only VESM3, and VESM++ (available at https://huggingface.co/ntranoslab/vesm).
VESM_3B and VESM3… See the full description on the dataset page: https://huggingface.co/datasets/ntranoslab/vesm_scores.dan_Fineweb2_edu_scoresjapanese-reranker-v2-hard-negatives-scores
hotchpotch/japanese-reranker-v2-hard-negatives-scores
This dataset was used to train the Japanese Reranker v2 family. It was not generated by Japanese Reranker v2 models.
Unified hard-negative score rows generated from the training sources used for Japanese Reranker v2-family experiments.
This dataset keeps teacher scores as floating-point soft labels. It follows the HPPRC Emb Score style: score/id rows and referenced document text collections are stored separately. Each score… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/japanese-reranker-v2-hard-negatives-scores.swe_Fineweb2_edu_scoresnob_Fineweb2_edu_scores
