datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
state-cancer-profiles
United States State Cancer Profiles data extract (mirror)
This is a mirror. Cite the Zenodo record, not this page:
Davis S. United States State Cancer Profiles data extract — vintage V3. Zenodo. https://doi.org/10.5281/zenodo.22085273
Concept DOI (always resolves to the latest vintage): https://doi.org/10.5281/zenodo.11098814
No DOI is minted on Hugging Face. HF hosts these bytes for native hf:// / DuckDB access and an ML audience that would never find the Zenodo record;… See the full description on the dataset page: https://huggingface.co/datasets/seandavis/state-cancer-profiles.laion-voice-profiles-annotated
LAION Voice Profiles — Annotated
Authors: Christoph Schuhmann and LAION.
28,212,933 utterances / 71,056 hours of synthetic English and German voice-acting speech from
500 distinct voice profiles, each driven through the same fixed matrix of 842 named acting
conditions. Every
utterance carries 40 emotion intensities, 57 perceptual voice dimensions, 4 audio-quality heads,
vocal-burst detections with timings, word-level forced alignment, MOSS audio codec tokens, a
768-d… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-annotated.laion-voice-profiles-dpo-cfg
LAION Voice Profiles — contrastive DPO pairs (CFG + phase 2)
Authors: Christoph Schuhmann and LAION.
842,935 preference pairs in four families, built from the same 500 synthetic voice profiles as
laion/laion-voice-profiles-sft
and laion/laion-voice-profiles-dpo.
These are the two pair families that the sister DPO set does not contain: they were built later,
for two measured defects of the models trained on it, and they are the complete remainder of the
project's preference… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-dpo-cfg.laion-voice-profiles-dpo
LAION Voice Profiles — TTS preference pairs (DPO)
Authors: Christoph Schuhmann and LAION.
3,451,531 preference pairs in three families, built from the same 500 synthetic voice profiles
as laion/laion-voice-profiles-sft.
Every pair shares one prompt; chosen and rejected differ only in the way the family names.
config
pairs
teaches
how rejected is made
emotion
1,064,594
hit the right tone for this line
a take of the same sentence, same voice, rendered at the wrong… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-dpo.moss-voice-profile-references
MOSS voice-profile references
Complete voice profiles: one reference speaker rendered through a matrix of named acting
conditions in English and German, with every candidate take kept — not just the winner — and
every take scored on itself.
Two releases live here.
voices
groups / voice
candidates / group
rows
audio
audio variants
pilot/ — the ten pilot voices
10
842
48
402,560
853.5 h
raw
root — Velvet Sage Baritone
1
832
32
106,424
239.6 h
raw, vc, raw_sidon… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/moss-voice-profile-references.recruitment-dataset-candidate-profiles-english
Djinni Dataset (English CVs part)
Overview
The Djinni Recruitment Dataset (English CVs part) contains 150,000 job descriptions and 230,000 anonymized candidate CVs, posted between 2020-2023 on the Djinni IT job platform. The dataset includes samples in English and Ukrainian.
The dataset contains various attributes related to candidate CVs, including position titles, candidate information, candidate highlights, job search preferences, job profile types, English… See the full description on the dataset page: https://huggingface.co/datasets/lang-uk/recruitment-dataset-candidate-profiles-english.laion-voice-profiles-sft
LAION Voice Profiles — TTS supervised fine-tuning set
Authors: Christoph Schuhmann and LAION.
1,200,531 instruction-tuning samples for reference-conditioned TTS, drawn from 500 synthetic
voice profiles: for each of the 842 acting conditions of each voice, the best 3 of its 48
candidate takes, each paired with a different clip of the same voice from another group as
the reference, plus the exact conditioning prompt, MOSS audio codes for both target and reference,
word-level… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-sft.profile-rollouts-v2mcp-sandbox-authority-boundary-profile
MCP Sandbox Authority Boundary Profile
Profile v0.1.0 · Release v0.2.0 - Experimental Characterization Profile
Profile release date: 2026-07-23
Latest distribution release date: 2026-09-05
Execution containment is not proof of bounded authority.
Start here
For a one-minute, case-by-case reading of the profile, open the companion
Authority Boundary Field Guide Space.
It presents the released synthetic observations with their control question,
observed result… See the full description on the dataset page: https://huggingface.co/datasets/msaleme/mcp-sandbox-authority-boundary-profile.eeyore_profilelinkedin-company-profilemazinger-dubber-profiles
Mazinger Dubber — Voice Profiles
Voice profiles for mazinger-dubber. Hosted on HuggingFace:
https://huggingface.co/datasets/bakrianoo/mazinger-dubber-profiles
Adding a New Profile
1. Prepare your files
Create a folder named after the profile:
profiles/
└── my-name/
├── script.txt # Plain-text transcript matching the audio exactly
└── voice.m4a # Voice sample (supported: .m4a, .wav, .mp3)
Tips: 10–30 seconds of clear speech, minimal… See the full description on the dataset page: https://huggingface.co/datasets/bakrianoo/mazinger-dubber-profiles.airep-embedded-evaluation-profile
AIREP Embedded Evaluation Profile v0.1
This is not a training dataset or benchmark. It is a Hugging Face distribution mirror of an
experimental evaluation-evidence profile, its schema basis, registry and fixtures. The canonical
specification history lives in the AIREP GitHub repository. Byte identity between this mirror and
its source commit is a distribution-integrity property, not independent scientific verification.
Experimental evaluation-evidence contract for AIREP v0.2.… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/airep-embedded-evaluation-profile.wan2.2-rocm-profiles
Wan2.2 Sequence Parallel ROCm Profiles (MI300X)
This dataset contains PyTorch/Perfetto traces, offline execution logs, serving benchmarks, and comparative reports for Wan2.2-T2V-A14B sequence-parallel runs on AMD ROCm (gfx942, 8x MI300X node).
Dataset Directory Structure
reports/ / Root:
wan22_rocm_sp_sweep_analysis.md: 3-way sequence parallel topology comparison report.
wan22_profile_u4_r1_analysis.md: Detailed analysis of the Ulysses-4 topology.… See the full description on the dataset page: https://huggingface.co/datasets/Akshat/wan2.2-rocm-profiles.task1192_food_flavor_profile
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1192_food_flavor_profile
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1192_food_flavor_profile.import-export-country-profile-v1
Global Import-Export Country Profile (1995-2024)
Dataset coverage and release
Coverage: 1995-2024 annual observations where available
Release version: 1.0.0
Stable repository identifier: yugalinks/import-export-country-profile-v1
Data format: Parquet
Zenodo archive: 10.5281/zenodo.22886794
License and use
Provider terms apply. Yugalinks prepares these files for analysis but does not relicense the underlying OECD observations or any CEPII reference… See the full description on the dataset page: https://huggingface.co/datasets/yugalinks/import-export-country-profile-v1.import-export-country-profile-series-v1
Global Import-Export Country Profile Series (1995-2024)
Dataset coverage and release
Coverage: 1995-2024 annual observations where available
Release version: 1.0.0
Stable repository identifier: yugalinks/import-export-country-profile-series-v1
Data format: Parquet
Zenodo archive: 10.5281/zenodo.22886791
License and use
Provider terms apply. Yugalinks prepares these files for analysis but does not relicense the underlying OECD observations or any… See the full description on the dataset page: https://huggingface.co/datasets/yugalinks/import-export-country-profile-series-v1.pythia-deduped-memorisation-profilesThis dataset has been created as an artefact of the paper Causal Estimation of Memorisation Profiles (Lesci et al., 2024).
More info about this dataset in the related collection Memorisation-Profiles.
import-export-hs2-profile-series-v1
Global Import-Export HS2 Profile Series (1995-2024)
Dataset coverage and release
Coverage: 1995-2024 annual observations where available
Release version: 1.0.0
Stable repository identifier: yugalinks/import-export-hs2-profile-series-v1
Data format: Parquet
Zenodo archive: 10.5281/zenodo.22886816
License and use
Provider terms apply. Yugalinks prepares these files for analysis but does not relicense the underlying OECD observations or any CEPII… See the full description on the dataset page: https://huggingface.co/datasets/yugalinks/import-export-hs2-profile-series-v1.import-export-route-profile-v1
Global Import-Export Route Profile (1995-2024)
Dataset coverage and release
Coverage: 1995-2024 annual observations where available
Release version: 1.0.0
Stable repository identifier: yugalinks/import-export-route-profile-v1
Data format: Parquet
Zenodo archive: 10.5281/zenodo.22886872
License and use
Provider terms apply. Yugalinks prepares these files for analysis but does not relicense the underlying OECD observations or any CEPII reference… See the full description on the dataset page: https://huggingface.co/datasets/yugalinks/import-export-route-profile-v1.import-export-product-profile-series-v1
Global Import-Export Product Profile Series (1995-2024)
Dataset coverage and release
Coverage: 1995-2024 annual observations where available
Release version: 1.0.0
Stable repository identifier: yugalinks/import-export-product-profile-series-v1
Data format: Parquet
Zenodo archive: 10.5281/zenodo.22886852
License and use
Provider terms apply. Yugalinks prepares these files for analysis but does not relicense the underlying OECD observations or any… See the full description on the dataset page: https://huggingface.co/datasets/yugalinks/import-export-product-profile-series-v1.import-export-hs4-profile-series-v1
Global Import-Export HS4 Profile Series (1995-2024)
Dataset coverage and release
Coverage: 1995-2024 annual observations where available
Release version: 1.0.0
Stable repository identifier: yugalinks/import-export-hs4-profile-series-v1
Data format: Parquet
Zenodo archive: 10.5281/zenodo.22886829
License and use
Provider terms apply. Yugalinks prepares these files for analysis but does not relicense the underlying OECD observations or any CEPII… See the full description on the dataset page: https://huggingface.co/datasets/yugalinks/import-export-hs4-profile-series-v1.africa-synth-energy-oilgas-field-profiles-nigeria
Africa Synth Energy Oilgas Field Profiles Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-oilgas-field-profiles-nigeria.personalaity-llm-personality-profiles
PersonalAIty: HEXACO personality profiles of frontier LLMs
Self-reported HEXACO personality profiles for 10 frontier language models across 8 vendors,
measured on 2026-08-16 with an open 50-item inventory, plus the instrument itself so the
measurement can be rerun or criticised.
This is a snapshot with a date on it, not a standing benchmark. Model versions drift; the
value here is that the whole measurement is reproducible with one command against models anyone
can reach.… See the full description on the dataset page: https://huggingface.co/datasets/Sciupy/personalaity-llm-personality-profiles.recruitment-dataset-candidate-profiles-ukrainian
Djinni Dataset (Ukrainian CVs part)
Overview
The Djinni Recruitment Dataset (Ukrainian CVs part) contains 150,000 job descriptions and 230,000 anonymized candidate CVs, posted between 2020-2023 on the Djinni IT job platform. The dataset includes samples in English and Ukrainian.
The dataset contains various attributes related to candidate CVs, including position titles, candidate information, candidate highlights, job search preferences, job profile types, English… See the full description on the dataset page: https://huggingface.co/datasets/lang-uk/recruitment-dataset-candidate-profiles-ukrainian.people-profiles-io-salesterraform-aws-ec2-instance-profile-codegenNIST-GenAI-Profile
NIST Generative AI Profile Question Answering Dataset
Dataset Summary
This dataset contains question-and-answer records derived from the National Institute of Standards and Technology publication NIST AI 600-1, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. The source publication is a cross-sectoral companion resource to the NIST AI Risk Management Framework and focuses on risks that are unique to, or exacerbated by… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/NIST-GenAI-Profile.iran_welfare_population_profile
ترکیب جمعیتی و توزیع دهکی — پایگاه رفاه ایرانیان
کدام استانها بیشترین سهم را در دهکهای پایین دارند، نرخ معلولیت و بیماری خاص کجا بالاتر است، و ساختار سنی و بُعد خانوار چطور تغییر میکند.
پوشش: 1402 · سطح: کشور و ۳۱ استان · تفکیک: دهک رفاهی × شهری/روستایی
تعداد مشاهده: 4,833 · تعداد شاخص: 6
منبع: وزارت تعاون، کار و رفاه اجتماعی — پایگاه رفاه ایرانیان
روش و محدودیتها
ارقام از نمونهٔ ۲ درصدیِ ناشناسِ پایگاه رفاه ایرانیان محاسبه شدهاند
(۱٬۶۹۷٬۸۱۶ فرد در… See the full description on the dataset page: https://huggingface.co/datasets/Farmaanaa/iran_welfare_population_profile.user_profile_dataset
User Profile Dataset
Public SFT dataset with separate Layer 1 and Layer 2 configurations.
User_Profile_L1: 69,000 train records and 1,952 test records
User_Profile_L2: 73,000 train records and 1,327 test records
Both configurations use deterministic random splits with seed 42. Layer 1
activities contain idx, Source, Action, and intent. Layer 1 records with
high-confidence grounding or duplication errors are removed before splitting.
