datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
proof-pileA dataset of high quality mathematical text.proofnetA dataset that evaluates formally proving and autoformalizing undergraduate mathematics.youtube-center-vietnamese-asrspot-terrain-dataset
Spot Dataset
nangang_sports_centervertebrate-v1-issue473-center1-cds
marin-dna/vertebrate-v1-issue473-center1-cds
Review status: draft generated for issue #473 review before upload.
Human-anchored 255 bp vertebrate sequences for the
cds cohort under the center_1 policy.
The policy projects the exact central human nucleotide, requires one target locus, and emits the 255 bp target window centered on that mapped nucleotide. Human anchors come from the fixed #417 protein-coding CDS anchor catalog.
The source projection was produced by the… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-center1-cds.details_Applied-Innovation-Center__AIC-1_v2
Dataset Card for Evaluation run of Applied-Innovation-Center/AIC-1
Dataset automatically created during the evaluation run of model Applied-Innovation-Center/AIC-1.
The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Applied-Innovation-Center__AIC-1_v2.vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered
marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered
Review status: draft generated for issue #473 review before upload.
Human-anchored 255 bp vertebrate sequences for the
ccre_enhancer_centered cohort under the full_window policy.
The policy projects the complete 255 bp human window, applies the established 128--512 bp compatible-fragment gate, and resizes around the accepted target-span midpoint. Human anchors come from the fixed exp351 ENCODE dELS/pELS-centered… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered.details_Applied-Innovation-Center__Karnak_v2
Dataset Card for Evaluation run of Applied-Innovation-Center/Karnak
Dataset automatically created during the evaluation run of model Applied-Innovation-Center/Karnak.
The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Applied-Innovation-Center__Karnak_v2.vertebrate-v1-issue473-center1-ccre-enhancer-centered
marin-dna/vertebrate-v1-issue473-center1-ccre-enhancer-centered
Review status: draft generated for issue #473 review before upload.
Human-anchored 255 bp vertebrate sequences for the
ccre_enhancer_centered cohort under the center_1 policy.
The policy projects the exact central human nucleotide, requires one target locus, and emits the 255 bp target window centered on that mapped nucleotide. Human anchors come from the fixed exp351 ENCODE dELS/pELS-centered catalog after… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-center1-ccre-enhancer-centered.zoonomia-v1-v4_ccre_enhancer_centered-order
bolinas-dna/zoonomia-v1-v4_ccre_enhancer_centered-order
An enhancer-CENTERED training set for issue
#351, built by the
snakemake/zoonomia_projection_dataset pipeline
(workflow/rules/centered.smk) at commit
8127acfea5aa.
Provenance
Each training window is defined directly from an ENCODE cCRE V4 enhancer
(dELS + pELS): one 255 bp window centered on the cCRE midpoint
(make_enhancer_anchors, keep-all — clustered enhancers each keep their own
window), rather than a… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_enhancer_centered-order.call-center-records
Call Center Records & Agent Performance Dataset (Free Sample)
This is a free sample with 2,013 rows. The full dataset has 19,959 rows across 3 tables.
Call detail records for a simulated insurance company call center with 25 agents
handling 35,000 calls over 12 months. Includes IVR menu paths, agent assignments,
call outcomes, hold times, transfer chains, and customer satisfaction scores.
Features realistic patterns: Monday morning surge, lunch dip, seasonal peaks
during open… See the full description on the dataset page: https://huggingface.co/datasets/mindweave/call-center-records.im3_open_source_data_center_atlas_v2026.02.09
IM3 Open Source Data Center Atlas v2026.02.09 — refined database
This repository preserves the IM3 Open Source Data Center Atlas v2026.02.09 and
adds a source-enriched, audited 43-column power-source table for all 1,479
source geometry records (1,474 unique IM3 IDs). The publication retains the exact
13 upstream columns plus 30 stable label, interpretation, and evidence fields.
Duplicate geometry records are intentionally retained.
Files… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/im3_open_source_data_center_atlas_v2026.02.09.vn-provinces-shopping-centers
Vietnam provinces shopping centers
Number of shopping centers / commercial centers as of 31 December. Coverage 2008-2024. Some provinces absent when zero. Year 2024 is preliminary. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison
Color key
Files
provinces… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-shopping-centers.iran_cpi_statistical_center
شاخص بهای مصرفکننده و تورم ایران — مرکز آمار ایران
نسخهٔ تکمنبعی از «شاخص بهای مصرفکننده و تورم» با دادههای مرکز آمار ایران.
پوشش: 1390-01 → 1405-03 · سطح: ملی · تعداد شاخص: 4
اگر میخواهید هر دو مرجع را یکجا و با نمای یکپارچه ببینید، از نسخهٔ کاملِ چندمرجعی استفاده کنید:
Farmaanaa/iran_cpi_and_inflation_multisource.
منبع: مرکز آمار ایران
دربارهٔ فرمانا
این مجموعهداده بخشی از بانک دادهٔ فرمانا است: گردآوری، پاکسازی و استانداردسازیِ
سریهای زمانیِ… See the full description on the dataset page: https://huggingface.co/datasets/Farmaanaa/iran_cpi_statistical_center.sob-ft-finetune-ready
SOB-FT Finetune Ready
~100k source rows → ~152k chat SFT examples for fine-tuning a small language model on JSON extraction (generation) and JSON error detection / repair (correction), with prompts aligned to our SOB extraction and zero-shot repair benchmarks.
Derived from mariem123kfg/sob-ft-extract (multi-source structured extraction corpus, excluding original SOB benchmark rows). Errors were injected in-house, then rows were materialized into ready-to-train prompt/target… See the full description on the dataset page: https://huggingface.co/datasets/seneca-center/sob-ft-finetune-ready.aaronschlegel_austin-animal-center-shelter-outcomes-and
Austin Animal Center Shelter Outcomes
30,000 shelter animals
Dataset Info
Source: Kaggle
Original Size: 3.24 MB
Kaggle Downloads: 9,138
Files: 2
Files
aac_shelter_cat_outcome_eng.csv
aac_shelter_outcomes.csv
Mirrored from Kaggle
CASP16CASP (Critical Assessment of Structure Prediction) is a community wide experiment to determine and advance the state of the art in computational structural biology. Every two years, participants are invited to submit models for a set of macromolecules and macromolecular complexes (proteins, RNA, ligands) for which the experimental structures are not yet public. In the latest CASP round, CASP16 in 2024, nearly 100 research groups from around the world submitted more than 80,000 models on 100+… See the full description on the dataset page: https://huggingface.co/datasets/center4casp/CASP16.call_center_data_csviran_gdp_statistical_center
حسابهای ملی و تولید ناخالص داخلی ایران — مرکز آمار ایران
نسخهٔ تکمنبعی از «حسابهای ملی و تولید ناخالص داخلی» با دادههای مرکز آمار ایران.
پوشش: 1390 → 1403-Q1 · سطح: ملی · تعداد شاخص: 6
اگر میخواهید هر دو مرجع را یکجا و با نمای یکپارچه ببینید، از نسخهٔ کاملِ چندمرجعی استفاده کنید:
Farmaanaa/iran_gdp_national_accounts.
منبع: مرکز آمار ایران
دربارهٔ فرمانا
این مجموعهداده بخشی از بانک دادهٔ فرمانا است: گردآوری، پاکسازی و استانداردسازیِ
سریهای زمانیِ… See the full description on the dataset page: https://huggingface.co/datasets/Farmaanaa/iran_gdp_statistical_center.Text-to-Tsqlkhayyam-challengeadaption-louisville-data-center-docs-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-louisville_data_center_docs (augmented)
This dataset contains planning commission staff reports, zoning code excerpts, and news transcripts regarding hyperscale data center developments in Louisville, Kentucky. The documents detail specific project proposals, such as the Camp Ground Road facility, including technical reviews on traffic, water usage, and environmental impact.… See the full description on the dataset page: https://huggingface.co/datasets/JaySmith502/adaption-louisville-data-center-docs-augmented.doom-defend-the-center-100k-oracle
Doom Defend the Center 100K oracle
100K rows dataset from 317 episodes of Doom Defend the Center scenario.
A programmatic oracle plays the scenario with access to privileged ViZDoom info: spin one way, freeze and fire when an enemy crosses the crosshair.
The dataset does not contain traces of privileged info and can be learned from RGB only.
Each row carries
frames.u8: one 100x160 RGB frame at time t.
labels.npz: turn (left/none/right) and shoot (yes/no) + actions at time t-2… See the full description on the dataset page: https://huggingface.co/datasets/anakin87/doom-defend-the-center-100k-oracle.pokemon-card-centering-measurements
Pokémon Card Centering Measurements — 320 eBay Listings, PSA-Window Pass Rates (2026)
Measured front-centering data for 320 Pokémon card eBay listing photos across 302 distinct cards, computer-vision-measured (corner detection + perspective correction, not estimated): left/right and top/bottom border share for each. 48.1% clear PSA's Gem Mint 10 centering window (worst side within 55/45); 89.7% clear the Mint 9 window (60/40). Sampled from cards a buyer had already flagged as… See the full description on the dataset page: https://huggingface.co/datasets/rrhagentbiz/pokemon-card-centering-measurements.All-Dataset-ReposAll-Dataset-Repos-2026-28resultscall-center
Call Center
This is a collection for evaluation of LLMs on call center scenarios.
req2test
Req2Test
Req2Test is a large-scale, multi-language dataset for requirement-driven
unit test generation, built to enable knowledge distillation from large
teacher LLMs to small (2B–4B) student models. Depending on architecture,
precision, quantization, and runtime, such models may fit within approximately
8 GB of VRAM.
Each record pairs a natural-language functional requirement with eight
aligned code artifacts — an interface declaration and a unit test class in
each of four… See the full description on the dataset page: https://huggingface.co/datasets/centertest/req2test.
