datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fannie-mae-loan-performancesolverport-solver-performance
SolverPort Solver Performance Dataset
Solver benchmark results across 8 solvers and 10 optimization problem families.
Metrics per Run
Runtime (seconds)
Optimality gap (%)
Feasibility status
Time to first feasible solution
Best bound and gap improvement rate
Search speed
PAR10 penalty score
Oracle regret vs best solver
Solvers Profiled
CP-SAT, HiGHS, CBC, SCIP, GLPK, Gurobi, MiniZinc, ALNS
Files
{instance_id}_performance.json — full… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/solverport-solver-performance.fannie-mae-loan-performance-rawstudent_performance
Student performance
The Student performance dataset from Kaggle.
Configuration
Task
Description
encoding
Encoding dictionary showing original values of encoded features.
math
Binary classification
Has the student passed the math exam?
writing
Binary classification
Has the student passed the writing exam?
reading
Binary classification
Has the student passed the reading exam?
Usage
from datasets importload_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/mstz/student_performance.frontierco-solver-performance
frontierco-solver-performance
Solver performance benchmark dataset produced by the FrontierCO Solver Arena.
Source
Built on profiles aligned with CO-Bench/FrontierCO.
Contents
Field
Description
instance_id
Unique instance identifier
problem_type
One of 8 CO problems
size
small / medium / large
difficulty
easy / hard
time_budget_sec
10 / 30 / 60 / 300
solver_id
One of 13 solvers
optimality_gap_pct
Gap to known optimum… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/frontierco-solver-performance.Performance-Marketing-Data
Performance Marketing Expert Dataset
Dataset Description
This dataset contains comprehensive performance marketing knowledge and logical reasoning patterns for Meta (Facebook/Instagram), Google Ads, and TikTok advertising platforms. It's designed for fine-tuning language models to understand brand verticals, performance marketing strategies, and develop reasoning capacity for creating winning ad campaigns.
Dataset Structure
Each example follows an… See the full description on the dataset page: https://huggingface.co/datasets/Sri-Vigneshwar-DJ/Performance-Marketing-Data.korean-embedding-performance-v1-performance-1m
Korean Embedding Performance v1 — 1M
Qwen3-Embedding-8B의 한국어 retrieval data-scale 실험을 위한 정확히
1,000,000-row 연구·비상업 contrastive dataset이다. release_eligible: false, 통합
라이선스 other이며 upstream source 조건을 재허가하지 않는다.
구성
계열
Rows
비율
역할
nlpai-lab/ko-triplet-v1.0
600,254
60.03%
넓은 한국어 QA/retrieval core
F2 Korean QA/instruction
287,000
28.70%
webfaq, mqa, koalpaca, realQA, komagpie
F2 retrieval task train-family
4,146
0.41%
MIRACL, MrTidy, MLDR
F2 PAWS-X… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-performance-1m.korean-embedding-performance-v1-ablation-200k
Korean Embedding Performance v1 — Ablation 200K
Qwen3-Embedding-8B의 한국어 retrieval continued fine-tuning에서 LoRA/DoRA/부분 및
full fine-tuning, loss, hard-negative 전략을 비교하기 위한 200,000-row 연구·비상업
성능 데이터다. release_eligible: false이며 통합 라이선스는 other다. upstream
source별 조건을 재허가하지 않는다.
구성
계열
Rows
역할
nlpai-lab/ko-triplet-v1.0@1f5d72d
100,254
넓은 한국어 QA/retrieval core
F2 Korean QA/instruction
68,000
webfaq, mqa, koalpaca, realQA, komagpie
F2 retrieval task… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-ablation-200k.korean-embedding-performance-v1-pilot-50k
Korean Embedding Performance v1 — Pilot 50K
주의: 이 revision은 공개 benchmark 성능 후보 학습에 사용하면 안 된다.
사후 15-task exact text-hash 감사에서 평가 query 고유 hash 4개가 확인됐다.
파이프라인·최적화 진단과 contamination ablation에만 남기며, 교체본은
ablation-200k이다.
Qwen3-Embedding 계열의 한국어 retrieval 성능 실험을 위한 50,000-row 연구용
contrastive dataset이다. 각 row는 instruction-aware query, positive passage 1개,
hard/easy negative passage 1–7개를 ms-swift embedding message schema로 저장한다.
사용 조건과 공개 범위
이 저장소의 통합 라이선스는 other다.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-pilot-50k.student-performance
Student Performance Synthetic Dataset v2
A reproducible, explicitly structured synthetic high-school student dataset for machine-learning benchmarking, data-engineering tests, educational analytics method development, and controlled fairness experiments.
Scope and non-claims
The data are entirely synthetic. Generator parameters are designed for internal coherence and are not calibrated to a specific real school system, demographic population, causal effect, or… See the full description on the dataset page: https://huggingface.co/datasets/neuralsorcerer/student-performance.fbref_football_player_performance_2024-2025
FBref Football Player Performance Dataset (2024-2025 Season)
Dataset Description
This dataset contains comprehensive performance statistics for 2273 professional football players during the 2024-2025 season. Sourced from FBref, it includes both traditional metrics (goals, assists) and advanced analytics (xG, xAG, progressive actions) across top European leagues.
Curated by: FBref
License: Publicly available football statistics (check FBref terms for redistribution)… See the full description on the dataset page: https://huggingface.co/datasets/alaa1234ah/fbref_football_player_performance_2024-2025.b2b-digital-marketing-performance-benchmarks
B2B & Ecommerce Performance Marketing Benchmarks
Maintained and published by Datametrik — Performance Marketing and Growth Agency.
marketing-campaign-performance-200k
Marketing Campaign Performance Dataset (200k)
Mirror de Kaggle: Marketing Campaign Performance Dataset (manishabhatt22), 200.000 filas de campañas de marketing (verificado: rango de fechas 2021-01-01 a 2021-12-31, no dos años como dice la card de Kaggle). Subido aquí para poder cargarlo con datasets.load_dataset() sin credenciales de Kaggle.
Contenido
data/marketing_campaign_dataset.csv — 200.000 filas, ~27 MB, 16 columnas.
data/data_dictionary.md — descripción… See the full description on the dataset page: https://huggingface.co/datasets/federicomoreno/marketing-campaign-performance-200k.fbref_football_player_performance_2024-2025
FBref Football Player Performance Dataset (2024-2025 Season)
Dataset Description
This dataset contains comprehensive performance statistics for 2273 professional football players during the 2024-2025 season. Sourced from FBref, it includes both traditional metrics (goals, assists) and advanced analytics (xG, xAG, progressive actions) across top European leagues.
Curated by: FBref
License: Publicly available football statistics (check FBref terms for redistribution)… See the full description on the dataset page: https://huggingface.co/datasets/3zden/fbref_football_player_performance_2024-2025.korean-embedding-performance-v1-sionic-retrieval-train-family-4146
Korean Sionic Retrieval Train-Family 4,146
F2LLM-v2가 공개한 Korean MIRACL, MrTidy, MLDR train-family row만 1M
decontaminated curriculum에서 lossless 추출한 target-adaptation dataset이다. 공개
evaluation query는 포함하지 않으며 current-student HN7 mining 전의 source artifact다.
구성과 목적
source
rows
역할
f2_miracl_ko_train
700
MIRACL Korean retrieval train-family
f2_mrtidy_korean_train
1,200
MrTidy Korean train
f2_mldr_ko_train
2,246
MLDR Korean long-document train-family
합계
4… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-retrieval-train-family-4146.korean-embedding-performance-v1-sionic-autorag-100k
Korean Embedding — Sionic AutoRAG domain 100K
AutoRAG의 금융·상거래·법률 domain retrieval을 보강하기 위한 100,000-row
performance dataset이다. F2LLM-v2 collection의 영어 FIQA/Amazon/Banking77과 중국어
e-commerce/legal QA를 query/positive/negative contrastive schema로 묶었다.
사용 조건과 평가 노출
release_eligible: false인 performance/non-commercial 연구용 composite다. 통합
라이선스는 other이며 F2 collection의 Apache-2.0 표기가 개별 upstream 권리를
재허가하지 않는다.
AutoRAG evaluation repository, query, qrel, corpus는 loader 입력으로… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-autorag-100k.africa-synth-trade-sez-ftz-performance-all
African SEZ FTZ Performance | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-trade-sez-ftz-performance-all.africa-synth-chw-performance-supervision-all
CHW Performance & Supervision Dataset (iCCM, Stockouts, Motivation) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-chw-performance-supervision-all.korean-embedding-performance-v1-sionic-squad-train-60k
Korean Embedding — Sionic SQuAD train-family 60K
KorQuAD v1.0의 원본 train split만 질문→정답 문맥 retrieval 형식으로 변환한
60,000-row target-adaptation 데이터다. Sionic retrieval 9종 중
SQuADKorV1의 train-family 신호를 명시적으로 보강한다.
사용 조건과 점수 공개 방식
release_eligible: false인 performance/non-commercial 실험용 composite다. 이
저장소의 통합 라이선스는 other이며 upstream 권리를 재허가하지 않는다. Hub metadata는
KorQuAD source를 CC-BY-ND-4.0으로 표시하고, upstream dataset card 본문은
CC BY-ND 2.0 KR도 명시한다. 사용자는 원 source 조건을 직접 확인해야 한다.
이… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-squad-train-60k.korean-embedding-performance-v1-sionic-health-100k
Korean Embedding — Sionic health multilingual 100K
Qwen3-Embedding 계열의 한국어 PublicHealthQA와 multilingual medical retrieval을
보강하기 위한 100,000-row performance dataset이다. F2LLM-v2 collection의 영어 중심
medical QA/instruction/flashcard와 소량 중국어 WebMedQA를 query/positive/negative
contrastive schema로 묶었다.
사용 조건
release_eligible: false인 performance/non-commercial 연구용 composite다. 통합
라이선스 표기는 other이며 collection card의 Apache-2.0 표기가 각 upstream source의
권리·개인정보·의료 데이터 조건을 재허가하지 않는다.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-health-100k.bnpl-credit-performanceusace-waterway-lock-performance
USACE Waterway Lock Inventory — Free
Complete inventory available free.
The complete 234-row physical lock inventory is available free on Hugging Face.
Verified coverage: 26 states and 60 waterways; age reference year 2024.
Records: 234 lock records; 56 fields.
This repository contains the complete public CSV: 234 rows and 56 fields. No paid purchase is needed.
Limitations
This is a physical inventory, not traffic, throughput, delay or performance history.… See the full description on the dataset page: https://huggingface.co/datasets/claritystorm/usace-waterway-lock-performance.Sorting-Algorithms-Performance-Metrics
Sorting Algorithms Benchmark Dataset (Array Size: 1000)
A benchmark dataset comparing execution time, memory usage, and comparison counts of various sorting algorithms (Bubble Sort, Selection Sort, Insertion Sort, Merge Sort, Quick Sort, Heap Sort, Odd-Even Sort) on arrays of size 1000. Each algorithm was run 100 times with randomized inputs to ensure statistical significance.
Dataset Details
Columns
run: Trial number (1-100 per algorithm).
algorithm:… See the full description on the dataset page: https://huggingface.co/datasets/ismielabir/Sorting-Algorithms-Performance-Metrics.africa-synth-retail-and-ecommerce-email-marketing-performance-data-nigeria
Email Marketing Performance Data | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-retail-and-ecommerce-email-marketing-performance-data-nigeria.Government-Performance-and-Results-Act
Government Performance and Results Act of 1993 Corpus
Dataset Description
The Government Performance and Results Act of 1993 Corpus is a processed legal and public-administration dataset derived from the Government Performance and Results Act of 1993, commonly abbreviated as GPRA.
GPRA established a statutory framework for strategic planning, annual performance planning, program performance measurement, and performance reporting across the executive branch of… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/Government-Performance-and-Results-Act.performance-tiers
Performance Tiers
Models grouped by speed:
Ultra Fast (30+ t/s): 1 models
Fast (15-30 t/s): 9 models
Moderate (5-15 t/s): 16 models
Slow (<5 t/s): 5 models
🚀 dispatchAI
Student_Performanceafrica-synth-cement-brand-performance-all
Africa Synth Cement Brand Performance All | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: infrastructure_transport - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-cement-brand-performance-all.metasolver-fjssp-solver-performance
MetaSolver-FJSP Solver Performance
Benchmark results for the 6-solver FJSSP portfolio under 30-second makespan minimization budget.
Metrics
Runtime, makespan @5s, makespan @30s, optimality gap, PAR10, oracle regret, selector accuracy.
External Benchmark
CO-Bench/FrontierCO
License
Apache 2.0
transformersjs-performance-leaderboard-results-dev
