datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-edu-score-2
📚 FineWeb-Edu-score-2
1.3 trillion tokens of the finest educational data the 🌐 web has to offer
What is it?
📚 FineWeb-Edu dataset consists of 1.3T tokens (FineWeb-Edu) and 5.4T tokens of educational web pages filtered from 🍷 FineWeb dataset. This is the 5.4 trillion version.
Note: this version uses a lower educational score threshold = 2, which results in more documents, but lower quality compared to the 1.3T version. For more details check the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu-score-2.gpn-star-scores
GPN-Star genome-wide scores
Genome-wide, mutation-rate-calibrated GPN-Star constraint and variant scores for
eight score sets covering human, mouse, chicken, D. melanogaster,
C. elegans, and A. thaliana. Canonical scores are available as
chromosome-sharded Parquet. Hugging Face hosts 72 BigWigs; a multi-assembly
UCSC track hub references the 64 logo/LLR views.
Overview and quick links
Resource
Link
Files
Browse all dataset files
Default Dataset… See the full description on the dataset page: https://huggingface.co/datasets/songlab/gpn-star-scores.unit4-students-scoresOpenDataArena-scored-data-2603
OpenDataArena-scored-data-2603
This repository provides a scored SFT dataset collection currently featuring 63 high-quality instruction-following datasets with nearly 25 million samples. The core value lies in its 30-dimensional scoring: every sample has been evaluated on metrics such as IFD, PPL, Deita_Quality, and 27 others, enabling fine-grained data selection for filtering, curriculum learning, and mixture optimization.
Key features:
30 metrics per sample — From lexical… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/OpenDataArena-scored-data-2603.TM-DATA_quality_score_v1
Dataset Card for "TM-DATA_quality_score_v1"
Adding quality score v1 to Locutusque/TM-DATA
More Information needed
matrix-scores
Dataset Description
This dataset contains the output of the MATRIX pipeline — Every Cure's computational drug repurposing scoring system. It provides ML-generated treatment probability scores for ~39.5 million drug-disease pairs, covering ~1,800 drugs × ~22,000 diseases.
⚠️ Research use only. These scores are the output of a computational research pipeline and do not constitute medical advice, clinical recommendations, or endorsement of any drug for any use. All findings… See the full description on the dataset page: https://huggingface.co/datasets/everycure/matrix-scores.cosmopedia_quality_score_v2Adding quality score v2 to HuggingFaceTB/cosmopedia
saiga_scoredSFT dataset for the Saiga family of models collected from various sources.
starcoderdata-python-edu-lang-score
Dataset Card for Starcoder Data with Python Education and Language Scores
Dataset Summary
The starcoderdata-python-edu-lang-score dataset contains the Python subset of the starcoderdata dataset. It augments the existing Python subset with features that assess the educational quality of code and classify the language of code comments. This dataset was created for high-quality Python education and language-based training, with a primary focus on facilitating models that can… See the full description on the dataset page: https://huggingface.co/datasets/JanSchTech/starcoderdata-python-edu-lang-score.openwebtext_quality_score_v1
Dataset Card for "openwebtext_quality_score_v1"
Adding quality score v1 to Skylion007/openwebtext
More Information needed
reranker-scores
Reranker-Scores
既存の日本語検索・QAデータセットについて、データセット中のクエリに付与された正・負例の関連度を多言語・日本語reranker 5種類を用いてスコア付けしたデータセットです。クエリごとに200件程度の正・負例文書が付与されています (事例ごとに付与個数はバラバラなのでご注意ください)。score.pos.avg および score.neg.avg はそれぞれ正例・負例についての5種のrerankerの平均スコアになっています。
Short Name
Hub ID
bge
BAAI/bge-reranker-v2-m3
gte
Alibaba-NLP/gte-multilingual-reranker-base
ruri
cl-nagoya/ruri-reranker-large
ruriv3-preview
cl-nagoya/ruri-v3-reranker-310m-preview
ja-ce… See the full description on the dataset page: https://huggingface.co/datasets/hpprc/reranker-scores.deep-scores-v2
DeepScoresV2 — Complete
A HuggingFace-formatted mirror of the complete version of the
DeepScoresV2 dataset for music object detection.
Dataset description
DeepScoresV2 is a large-scale dataset of synthetically rendered music score pages
annotated with bounding boxes for musical symbols. The complete version contains
255,385 images with 151 million annotated instances across 135 symbol classes.
Each image is a full score page rendered from MuseScore across 5 music fonts… See the full description on the dataset page: https://huggingface.co/datasets/zzsi/deep-scores-v2.iDRAMA-scored-2024
Dataset Summary
iDRAMA-Scored-2024 is a large-scale dataset containing approximately 57 million social media posts from web communities on social media platform, Scored.
Scored serves as an alternative to Reddit, hosting banned fringe communities, for example, c/TheDonald, a prominent right-wing community, and c/GreatAwakening, a conspiratorial community.
This dataset contains 57M posts from over 950 communities collected over four years, and includes sentence embeddings for all… See the full description on the dataset page: https://huggingface.co/datasets/iDRAMALab/iDRAMA-scored-2024.stockimage-scored-pt12refinedweb-3m_quality_score_v1
Dataset Card for "refinedweb-3m_quality_score_v1"
Adding quality score v1 to mattymchen/refinedweb-3m
More Information needed
hpprc_emb_reranker_score
⚠️ お知らせ
よりスコア付したデータ件数とrerankerのバリエーションを増やしたデータセットのhotchpotch/hpprc_emb-scoresも公開しています。
hpprc/emb (便利なデータセットの公開、ありがとうございます)の collection と dataset がペアになっているデータに対し、negative を最大32個ランダムサンプリングしたものを、hotchpotch/japanese-bge-reranker-v2-m3-v1でスコア付けしたものです。
ライセンスは、subset ごとに hpprc/emb に記載のライセンスと同等とします。
スコア作成タイミングの revision に対してスコアを付与しているため、revision を変えると場合によって行ズレやデータ構造の変化が発生する可能性があることに注意が必要です。
例
from datasets import load_dataset
# targets = ("auto-wiki-qa", "4feb2e2492")… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/hpprc_emb_reranker_score.OpenDataArena-scored-data
OpenDataArena-scored-data
This repository contains a special scored and enhanced collection of over 47 original datasets for Supervised Fine-Tuning (SFT).
These datasets were processed using the OpenDataArena-Tool, a comprehensive suite of automated evaluation methods for assessing instruction-following datasets.
The scored-dataset collection from the OpenDataArena (ODA) reflects a large-scale, cross-domain effort to quantify dataset value in the era of large language models. ODA… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/OpenDataArena-scored-data.fineweb-edu-bottom1pct-scorestockimage-scored-pt6stockimage-scored-pt8fineweb-edu-bottom1pct-lang-scorestockimage-scored-pt1BIGstockimage-1.5M-scored-pt-twommarco-hard-negatives-reranker-score
hotchpotch/mmarco-hard-negatives-reranker-score
This repository contains data from mMARCO scored using the reranker BAAI/bge-reranker-v2-m3.
Languages Covered
target_languages = [
"english",
"chinese",
"french",
"german",
"indonesian",
"italian",
"portuguese",
"russian",
"spanish",
"arabic",
"dutch",
"hindi",
"japanese",
"vietnamese"
]
Hard Negative Data
The hard negative data is derived from… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/mmarco-hard-negatives-reranker-score.stockimage-scored-pt9fineweb-edu-top1pct-lang-scorefineweb-edu-top1pct-scorestockimage-scored-pt10BIGstockimage-1.5M-scored-pt-onepairs_with_scores_v27
