datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WITH_SCOREHIGH_SCOREpavlick-formality-scoresThis dataset contains sentence-level formality annotations used in the 2016
TACL paper "An Empirical Analysis of Formality in Online Communication"
(Pavlick and Tetreault, 2016). It includes sentences from four genres (news,
blogs, email, and QA forums), all annotated by humans on Amazon Mechanical
Turk. The news and blog data was collected by Shibamouli Lahiri, and we are
redistributing it here for the convenience of other researchers. We collected
the email and answers data ourselves, using… See the full description on the dataset page: https://huggingface.co/datasets/osyvokon/pavlick-formality-scores.NO_SCOREhumaneval-rerun-scoresArabic-NLi-Pair-Score
Arabic NLI Pair-Score
Dataset Summary
The Arabic Version of SNLI and MultiNLI datasets. (Pair-Score Subset)
Originally used for Natural Language Inference (NLI),
Dataset may be used for training/finetuning an embedding model for semantic textual similarity.
Pair-Class Subset
Columns: "sentence1", "sentence2", "score"
Column types: str, str, float
Arabic Examples:
{
"sentence1": "شخص على حصان يقفز فوق طائرة معطلة",
"sentence2": "شخص يقوم… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic-NLi-Pair-Score.resume-ats-score-v1-en
Resume-ATS Score Dataset v1 (English)
Dataset Description
resume-ats-score-v1-en is a semantic similarity dataset designed for training sentence transformers to predict ATS (Applicant Tracking System) compatibility scores between resumes and job descriptions. This dataset enables fine-tuning models to understand the semantic alignment and matching quality between candidate profiles and job requirements.
Key Features
📊 6.4K examples (5.1K train, 1.3K… See the full description on the dataset page: https://huggingface.co/datasets/0xnbk/resume-ats-score-v1-en.agentic-score-leaderboard
🛠️ Agentic Score Leaderboard — one RTX 5090
How well do local models actually drive a tool-using agent loop? Not single-call function-calling
benchmarks — a real loop: native OpenAI tool-calling through llama-server, multi-step deterministic
tasks, programmatic verification. Everything runs on a single RTX 5090 32GB.
Updated 2026-06-17 · llama.cpp b9562 · --jinja native tool-calling · temp 0.
Leaderboard
#
model
params
Agentic Score
success
tool-eff… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/agentic-score-leaderboard.COMET_scoremturk_scorestdist-msmarco-scores
MS MARCO Distillation Scores for Translate-Distill
This repository contains MS MARCO training
query-passage scores produced by MonoT5 reranker
unicamp-dl/mt5-13b-mmarco-100k and
castorini/monot5-3b-msmarco-10k.
Each training query is associated with the top-50 passages retrieved by the ColBERTv2 model.
Files are gzip compressed and with the naming scheme of {teacher}-monot5-{msmarco, mmarco}-{qlang}{plang}.jsonl.gz,
which indicates the teacher reranker that inferenced using… See the full description on the dataset page: https://huggingface.co/datasets/hltcoe/tdist-msmarco-scores.release-scores
TuneJury Reward Scores
Pre-computed TuneJury reward scores for seven open-license music collections (219,020 clips total). Companion artifact to the paper TuneJury: An Open Metric for Improving Music Generation Preference Alignment (arXiv:2606.17006, code).
This dataset ships scores and identifiers only, not audio. Each row is one deterministic TuneJury scorer call per clip. To obtain the audio, fetch each collection from its original source (see "Audio sources" below).
The… See the full description on the dataset page: https://huggingface.co/datasets/TuneJury/release-scores.polymarket-kalshi-scoresync-orderbook-sample
Polymarket x Kalshi Orderbook Archive - Free Score-Synced Sample
A free excerpt of the ZenHodl Polymarket & Kalshi Historical Orderbook Archive,
published so you can VERIFY the two non-reconstructable features before you buy --
the things competitors (Telonex, PolymarketData) do not ship:
Per-row score-sync - every Kalshi row carries the live game state next to the quote:
home_score, away_score, score_diff, period, time_remaining, game_state.
Cross-venue - the same games are… See the full description on the dataset page: https://huggingface.co/datasets/Coyevans/polymarket-kalshi-scoresync-orderbook-sample.ceo-pay-scorecard
CEO Pay-vs-Delivery Scorecard: S&P 500 v2.2
A frozen, reproducible dataset for a descriptive audit of granted compensation, Compensation Actually Paid and shareholder return across a dated S&P 500 universe.
AI disclosure: the research is the author's; this text was drafted with AI assistance and reviewed by the author. The model, and the conflict it creates, are named in the Conflict of interest section below.
Author: NM AI Research, ORCID 0009-0003-4213-7769
Canonical concept… See the full description on the dataset page: https://huggingface.co/datasets/NMAIResearch/ceo-pay-scorecard.agent-discoverability-ado-score-romania
Agent Discoverability (ADO Score) — Romania, September 2026
130 Romanian domains probed for A2A Agent Cards, MCP discovery, llms.txt, schema.org and Wikidata. Zero Agent Cards; mean ADO Score 17/100. Raw data, scripts and scoring spec, CC BY 4.0.
Canonical study (analysis, charts, interpretation):
Romanian ·
English
What this is
On 8 September 2026 a standard-library Python probe (published) requested, for each of 130 domains, the homepage without JavaScript… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/agent-discoverability-ado-score-romania.vhh_affinity-score
Nanobody (VHH) Affinity Prediction Dataset
Dataset Overview
This dataset helps predict the binding affinity between nanobodies (VHH, single-domain antibodies from camelids) and their target antigens. Affinity is a key parameter that measures how strongly an antibody binds to its antigen, usually expressed as dissociation constant (KD) or binding free energy.
High affinity is a critical property for therapeutic antibodies, so accurately predicting nanobody affinity is… See the full description on the dataset page: https://huggingface.co/datasets/ZYMScott/vhh_affinity-score.empathy_scoresclust_pathway_scoresnfl_bets_scoressentiment-headline-scores
Sentiment Headline Scores
91,851 labeled news headlines for S&P/DOW/NASDAQ stocks (Reuters/Eikon, July 2019 to Oct 2020), scored using a lexicon I built for a course homework: ShubhamOza/sentiment-headline-lexicon.
Columns
column
what it is
ticker
stock ticker the headline is about
time
headline date
headlines
raw headline text
returns
next period return
label
1.0 if the return was positive, -1.0 if negative
PARTITION_SAMPLE
train / test /… See the full description on the dataset page: https://huggingface.co/datasets/ShubhamOza/sentiment-headline-scores.mt_bench_single_score_gpt4_judgementskin-score-322-foods
SpotBite Skin Score Dataset — 322 foods scored for acne-related diet factors
A small, open dataset of 322 common foods, each scored 0–100 for how likely one typical serving is to contribute to acne, based on four diet factors that published trials link to breakouts: glycemic load, dairy, inflammatory fats, and anti-inflammatory boosters (omega-3, polyphenols).
The scores are the same ones shown in the SpotBite app and on its public food pages (spotbite.app/foods/<slug>). Every… See the full description on the dataset page: https://huggingface.co/datasets/SpotBite/skin-score-322-foods.kitgrade-scores
KitGrade: the component scores behind every kit grade
One row per kit carrying the six component scores the total is composed from, the method version that produced them, and the full breakdown JSON.
Rows in this cut
36
One row is
one kit
Cut
2026-09-04
Refreshed
Monthly, on the first of the month
Measured by
KitGrade
Method
https://toolproof.thecompound.tech/methodology
Licence
Creative Commons Attribution 4.0 International
Publisher
Compound Labs… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/kitgrade-scores.resume-ats-score-v1-en
Resume-ATS Score Dataset v1 (English)
Dataset Description
resume-ats-score-v1-en is a semantic similarity dataset designed for training sentence transformers to predict ATS (Applicant Tracking System) compatibility scores between resumes and job descriptions. This dataset enables fine-tuning models to understand the semantic alignment and matching quality between candidate profiles and job requirements.
Key Features
📊 6.4K examples (5.1K train… See the full description on the dataset page: https://huggingface.co/datasets/sachanshreyas/resume-ats-score-v1-en.sleep-score-fitbit
Fitbit Sleep Score Data
About the Dataset
Description
The Fitbit Sleep Score dataset, available on Kaggle, comprises detailed sleep data sourced from an individual's Fitbit device. It includes metrics such as overall sleep score, revitalization score, deep sleep duration, resting heart rate, and restlessness, each timestamped for in-depth analysis.
Data Fields
timestamp: The specific date and time the sleep data was recorded.
overall_score: An… See the full description on the dataset page: https://huggingface.co/datasets/aai530-group6/sleep-score-fitbit.videophy_autoeval_scoresProject github: https://github.com/Hritikbansal/videophy
Paper: https://arxiv.org/abs/2406.03520
These scores are calculated using our auto-evaluator (https://huggingface.co/videophysics/videocon_physics/tree/main) on the test data (https://huggingface.co/datasets/videophysics/videophy_test_public).
Gemini_3.1_202_Task_AI_Exposure_Scores
Gemini 3.1 2026 Task AI Exposure Scores
Dataset Summary
This dataset contains task-level AI exposure labels for O*NET task statements. Each task is classified into one of four categories, E0, E1, E2, or E3, using an updated 2026 Agentic AI Exposure Rubric and a Gemini 3.1 Pro classification pipeline. The labels are designed to capture whether a task can be accelerated by a frontier agentic AI system directly, whether it would require deeper software integration, or… See the full description on the dataset page: https://huggingface.co/datasets/MIT-WAL/Gemini_3.1_202_Task_AI_Exposure_Scores.MMLU-Pro-reasoning-score
Dataset Card for MMLU Pro with reasoning scores
MMLU Pro dataset with reasoning scores
Dataset Details
Dataset Description
As discovered in "When an LLM is apprehensive about its answers -- and when its uncertainty is justified", amount of reasoning required to answer a question (a.k.a. reasoning score) is a beter metric to estimate model uncertainty compared to more human-like level of education. Following the foot steps outlined in that paper, we ask a… See the full description on the dataset page: https://huggingface.co/datasets/LabARSS/MMLU-Pro-reasoning-score.korea-housing-subscription-score-2026
2026 Korea Private Housing Subscription Score Table
A reusable CSV dataset for Korea's private-housing subscription point system.
The maximum total score is 84 points:
No-home period: up to 32 points
Dependents: up to 35 points
Housing-subscription-account duration: up to 17 points
Spouse account-duration recognition can add up to 3 points, while the combined account-duration category remains capped at 17 points.
Original source and methodology… See the full description on the dataset page: https://huggingface.co/datasets/eunguneun/korea-housing-subscription-score-2026.resume-ats-score-v1-en
Resume-ATS Score Dataset v1 (English)
Dataset Description
resume-ats-score-v1-en is a semantic similarity dataset designed for training sentence transformers to predict ATS (Applicant Tracking System) compatibility scores between resumes and job descriptions. This dataset enables fine-tuning models to understand the semantic alignment and matching quality between candidate profiles and job requirements.
Key Features
📊 6.4K examples (5.1K train… See the full description on the dataset page: https://huggingface.co/datasets/Prakhar141205/resume-ats-score-v1-en.
