datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dParallel_LLaDA_Distill_Data
dParallel-LLaDA-Distill Dataset:
This dataset is used for the certainty-forcing distillation process in dParallel. We use prompts from publicly available training datasets and let the pretrained model generate its own responses as training data. For LLaDA-8B-Instruct, we sample prompts from the GSM8K, PRM12K training set, and part of the Numina-Math dataset. We generate target trajectories using a semi-autoregressive strategy with a sequence length of 256 and block length of 32. We… See the full description on the dataset page: https://huggingface.co/datasets/Zigeng/dParallel_LLaDA_Distill_Data.dpa-2-to-4-unit-2026
Dataset Card for Down Payment Assistance Survey (83 US Metros, 2026)
Dataset Details
Homepage / methodology: https://vantovault.com/library/down-payment-assistance-duplex/
Maintained at: Van to Vault (vantovault.com)
License: CC BY 4.0 — use with attribution to Van to Vault (vantovault.com)
Point of contact: Stephan D., Van to Vault
Dataset Summary
A hand-checked survey answering one narrow question across 83 U.S. metro areas: can an
owner-occupant… See the full description on the dataset page: https://huggingface.co/datasets/vantovault/dpa-2-to-4-unit-2026.WildChat-2k-TypeTopic
WildChat-2k-TypeTopic: The Manually Curated Edition
Dataset Description
WildChat-2k-TypeTopic is a manually curated subset of 1,880 real-world user prompts from the WildChat dataset, featuring annotations for both task type (e.g. knowledge recall, problem solving, creative, lists) and topic category (e.g. personal assistance, math, ai, household)
Why this dataset?
Suppose you want to answer a research question such as "What kind of user prompt does the LLM like… See the full description on the dataset page: https://huggingface.co/datasets/dpaleka/WildChat-2k-TypeTopic.dParallel_Dream_Distill_Data
dParallel-Dream-Distill Dataset:
This dataset is used for the certainty-forcing distillation process in dParallel. We use prompts from publicly available training datasets and let the pretrained model generate its own responses as training data. For LLaDA-8B-Instruct, we sample prompts from the GSM8K, PRM12K training set, and part of the Numina-Math dataset. We generate target trajectories using a semi-autoregressive strategy with a sequence length of 256 and block length of 32. We… See the full description on the dataset page: https://huggingface.co/datasets/Zigeng/dParallel_Dream_Distill_Data.DPAO_filterroof-segmentation-control-netdcd_pairwise_setup_basedpa-programs-2026
DPA Programs 2026
34 down payment assistance programs across 13 states.
Details
Records: 34
Format: JSONL
License: CC-BY-4.0
Last Updated: March 2026
Verified By: Beau Thompson, NMLS #1615561
Publisher: Good News Lending
Thompson Alpha Logic
State-by-state mapping of 2026 DPA grants cross-referenced with FHA/USDA eligibility to find 'Net-Zero Cash' purchase zones where buyers can close with $0 out of pocket by stacking DPA with zero-down loan programs.
This… See the full description on the dataset page: https://huggingface.co/datasets/Good-News-Lending/dpa-programs-2026.s2-sr-control-netroof3d-segmentation-control-netamazon-image-title-triplet-250k-cleanedDPATeX-CSV-1.0Chaalan et al., DPATeX2025: a dataset for DPA Mathematical Perturbation Core.
dpa-programs-2026
Down Payment Assistance Programs 2026
Comprehensive database of down payment assistance (DPA) programs across 13 Southeast US states, curated by Good News Lending.
Dataset Description
34 active DPA programs from state Housing Finance Agencies (HFAs), including grants, forgivable loans, deferred loans, and repayable second mortgages.
States Covered (13)
Tennessee, Mississippi, Alabama, Georgia, Florida, Kentucky, Louisiana, North Carolina, South Carolina… See the full description on the dataset page: https://huggingface.co/datasets/wendymthompson/dpa-programs-2026.amazon-image-title-tripletd0CIFAR10FakeKaggleamazon-image-title-triplet-250kteacher-assistant-instruction-datasetkl3m-filter-data-dotgov-www.dpaa.milLing-Coder-dParallel-merged-512-120kFB-email-market
Dataset Card for "FB-email-market"
More Information needed
kl3m-data-dotgov-www.dpaa.mil
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.dpaa.mil.descripciones-frutas-y-verdurasEsta base de datos consiste diferentes descripciones de hasta (en su última revisión) 10 alimentos (todas ellas frutas y verduras).
La finalidad de esta base de datos es realizar fine-tuning al modelo gplsi/Aitana-2B-S-base para el proyecto Capstone-project Samsung AI Innovation Campus
Personalization_Bench_w_dpapp-progressive-leftauguste_pocplanetarian-dpa-stage1-nojp
planetarian-dpa-stage1-nojp
This dataset was created using the Claude Dataset Skill.
personalization_prompt_response_mistral_dpatest_mistral_dpa
