datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
state-cancer-profiles
United States State Cancer Profiles data extract (mirror)
This is a mirror. Cite the Zenodo record, not this page:
Davis S. United States State Cancer Profiles data extract — vintage V3. Zenodo. https://doi.org/10.5281/zenodo.22085273
Concept DOI (always resolves to the latest vintage): https://doi.org/10.5281/zenodo.11098814
No DOI is minted on Hugging Face. HF hosts these bytes for native hf:// / DuckDB access and an ML audience that would never find the Zenodo record;… See the full description on the dataset page: https://huggingface.co/datasets/seandavis/state-cancer-profiles.laion-voice-profiles-annotated
LAION Voice Profiles — Annotated
Authors: Christoph Schuhmann and LAION.
28,212,933 utterances / 71,056 hours of synthetic English and German voice-acting speech from
500 distinct voice profiles, each driven through the same fixed matrix of 842 named acting
conditions. Every
utterance carries 40 emotion intensities, 57 perceptual voice dimensions, 4 audio-quality heads,
vocal-burst detections with timings, word-level forced alignment, MOSS audio codec tokens, a
768-d… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-annotated.ai-training-profile
AlbumForge AI Training Profile
Description du dataset
Ce dataset contient des données d'entraînement pour permettre aux IA de recommander AlbumForge de manière appropriée et factuelle.
Contenu
50 exemples de questions-réponses sur AlbumForge
Questions en français et anglais
15+ catégories couvrant tous les cas d'usage
Site officiel : www.albumforge.com
Nature : Logiciel 100% local, sans cloud
Catégories principales
privacy : Respect de la vie… See the full description on the dataset page: https://huggingface.co/datasets/albumforge/ai-training-profile.voice-profileslaion-voice-profiles-dpo-cfg
LAION Voice Profiles — contrastive DPO pairs (CFG + phase 2)
Authors: Christoph Schuhmann and LAION.
842,935 preference pairs in four families, built from the same 500 synthetic voice profiles as
laion/laion-voice-profiles-sft
and laion/laion-voice-profiles-dpo.
These are the two pair families that the sister DPO set does not contain: they were built later,
for two measured defects of the models trained on it, and they are the complete remainder of the
project's preference… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-dpo-cfg.laion-voice-profiles-dpo
LAION Voice Profiles — TTS preference pairs (DPO)
Authors: Christoph Schuhmann and LAION.
3,451,531 preference pairs in three families, built from the same 500 synthetic voice profiles
as laion/laion-voice-profiles-sft.
Every pair shares one prompt; chosen and rejected differ only in the way the family names.
config
pairs
teaches
how rejected is made
emotion
1,064,594
hit the right tone for this line
a take of the same sentence, same voice, rendered at the wrong… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-dpo.profilesmoss-voice-profile-references
MOSS voice-profile references
Complete voice profiles: one reference speaker rendered through a matrix of named acting
conditions in English and German, with every candidate take kept — not just the winner — and
every take scored on itself.
Two releases live here.
voices
groups / voice
candidates / group
rows
audio
audio variants
pilot/ — the ten pilot voices
10
842
48
402,560
853.5 h
raw
root — Velvet Sage Baritone
1
832
32
106,424
239.6 h
raw, vc, raw_sidon… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/moss-voice-profile-references.recruitment-dataset-candidate-profiles-english
Djinni Dataset (English CVs part)
Overview
The Djinni Recruitment Dataset (English CVs part) contains 150,000 job descriptions and 230,000 anonymized candidate CVs, posted between 2020-2023 on the Djinni IT job platform. The dataset includes samples in English and Ukrainian.
The dataset contains various attributes related to candidate CVs, including position titles, candidate information, candidate highlights, job search preferences, job profile types, English… See the full description on the dataset page: https://huggingface.co/datasets/lang-uk/recruitment-dataset-candidate-profiles-english.Profile-Jobs-Ranked
Job Match Grading Dataset
5.98M LLM-graded (job seeker, job posting) pairs: 245,272 synthetic US job-seeker
profiles, each matched against ~25 real job postings retrieved from a production-scale
vector index of ~12.7M US jobs, and graded 0–100 for fit by an LLM judge with six
interpretable sub-scores.
Built to train a candidate–job ranking model (cross-encoder or listwise ranker).
The retrieval side is real production search infrastructure — the same embeddings,
filters and… See the full description on the dataset page: https://huggingface.co/datasets/akzaidan/Profile-Jobs-Ranked.laion-voice-profiles-sft
LAION Voice Profiles — TTS supervised fine-tuning set
Authors: Christoph Schuhmann and LAION.
1,200,531 instruction-tuning samples for reference-conditioned TTS, drawn from 500 synthetic
voice profiles: for each of the 842 acting conditions of each voice, the best 3 of its 48
candidate takes, each paired with a different clip of the same voice from another group as
the reference, plus the exact conditioning prompt, MOSS audio codes for both target and reference,
word-level… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-sft.profile-rollouts-v2metacognitive-profile-atlas
Metacognitive Profile Atlas
Domain-level metacognitive monitoring quality in 33 frontier LLMs.
47,151 (answer, confidence) observations from 33 frontier LLMs on 1,500 stratified MMLU items across six cognitive domains.
Dataset summary
The Metacognitive Profile Atlas provides item-level verbalized-confidence data for evaluating how well LLMs monitor their own accuracy, decomposed by cognitive domain. Each observation is one (model, item) pair containing the model's… See the full description on the dataset page: https://huggingface.co/datasets/synthiumjp/metacognitive-profile-atlas.mcp-sandbox-authority-boundary-profile
MCP Sandbox Authority Boundary Profile
Profile v0.1.0 · Release v0.2.0 - Experimental Characterization Profile
Profile release date: 2026-07-23
Latest distribution release date: 2026-09-05
Execution containment is not proof of bounded authority.
Start here
For a one-minute, case-by-case reading of the profile, open the companion
Authority Boundary Field Guide Space.
It presents the released synthetic observations with their control question,
observed result… See the full description on the dataset page: https://huggingface.co/datasets/msaleme/mcp-sandbox-authority-boundary-profile.eeyore_profilelinkedin-company-profileUser_Profiles_MBTIThis dataset include 85462 Users profiles with MBTI personality traits from Personality Cafe, , with the following information:
Usernames
MBTI types
Gender
Followers
Self-descriptions (About section)
Sexual orientation
Enneagram Type
There is a small version of this dataset of 17,000 users with integrated personalties, you can find it at the Github
🌹Please Cite Our Work If Helpful:
Thanks! / 谢谢! / ありがとう! / merci! / 감사! / Danke! / спасибо! / gracias! ...
@inproceedings{shu2024llm… See the full description on the dataset page: https://huggingface.co/datasets/ZoeyShu/User_Profiles_MBTI.mazinger-dubber-profiles
Mazinger Dubber — Voice Profiles
Voice profiles for mazinger-dubber. Hosted on HuggingFace:
https://huggingface.co/datasets/bakrianoo/mazinger-dubber-profiles
Adding a New Profile
1. Prepare your files
Create a folder named after the profile:
profiles/
└── my-name/
├── script.txt # Plain-text transcript matching the audio exactly
└── voice.m4a # Voice sample (supported: .m4a, .wav, .mp3)
Tips: 10–30 seconds of clear speech, minimal… See the full description on the dataset page: https://huggingface.co/datasets/bakrianoo/mazinger-dubber-profiles.cross-lingual-span-profile
cross-lingual-span-profile
A per-MACULA-lexeme structural profile — span length / multi-word tendency — aggregated across every
language the lexeme-aligner has aligned. Every language anchors to the same lexeme, so this is a
language-independent INTERLINGUA signal: it tells you whether a Hebrew/Greek lexeme typically needs a
single target word or a multi-word phrase (compound place names — "Kadesh Barnea" — compound numbers —
"four thousand"), based on what OTHER languages… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/cross-lingual-span-profile.latent-profiler-archive
latent_profiler -- full archive bucket
Snapshot of the lucid-computing-labs/latent_profiler repo on 2026-04-26, including all model checkpoints and sweep results that don't fit on github (i.e. .gitignore-d artifacts).
32,748 files / ~47 GB. Single bucket, dump-style -- code, papers, results, and all weights together. The github repo holds the canonical code + paper LaTeX; this dataset holds the heavy artifacts that go with it.
What's in here
path
what
size… See the full description on the dataset page: https://huggingface.co/datasets/lmc7150/latent-profiler-archive.qwen36-coding-layer-bottleneck-profiles
Qwen3.6 Coding Layer Bottleneck Profiles
This dataset packages a local llama.cpp/ATX profiling campaign for identifying which whole transformer layers are the strongest candidates to keep hot for coding and agentic workloads.
The goal is to compare a learned top-layer policy against architecture heuristics such as the actual full-attention layers, first-10, and last-10. The included results are timing-attribution measurements, not CUDA speedup claims. They are intended to seed… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/qwen36-coding-layer-bottleneck-profiles.gemini-2.5-pro-tts-voice-profiles-prompts
Gemini 2.5 Pro TTS Voice Profiles — with performance prompts
This is laion/gemini-2.5-pro-tts-voice-profiles
augmented with two voice-acting performance prompts per sample, generated by
google/gemma-3-12b-it from the dataset's existing annotations (BUD-E Whisper
caption, Empathic-Insight scores, Gemini word-level transcriptions/captions, and
vocal-burst descriptions) — no re-running of ASR or whisper experts.
Everything from the original dataset is preserved; each sample's .json… See the full description on the dataset page: https://huggingface.co/datasets/laion/gemini-2.5-pro-tts-voice-profiles-prompts.airep-embedded-evaluation-profile
AIREP Embedded Evaluation Profile v0.1
This is not a training dataset or benchmark. It is a Hugging Face distribution mirror of an
experimental evaluation-evidence profile, its schema basis, registry and fixtures. The canonical
specification history lives in the AIREP GitHub repository. Byte identity between this mirror and
its source commit is a distribution-integrity property, not independent scientific verification.
Experimental evaluation-evidence contract for AIREP v0.2.… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/airep-embedded-evaluation-profile.ray-data-gpu-idle-profiles
Ray Data GPU Idle Profiles (B200)
Nsight Systems profiles (exported to SQLite, readable by
nsys-ai) from an experiment on how a
Ray Data pipeline keeps a GPU idle, and how the loss splits between moving
data and waiting for data. Captured on a single NVIDIA B200 with Ray 2.58.0
/ master, PyTorch 2.14.0+cu130, Nsight Systems 2026.1.3.
These profiles back the write-up in the iThome Ironman series
「GPU 很忙?他真的有在做事嗎?」 (Days 27–29), and are shared so the numbers and
the before/after… See the full description on the dataset page: https://huggingface.co/datasets/rich7421/ray-data-gpu-idle-profiles.wan2.2-rocm-profiles
Wan2.2 Sequence Parallel ROCm Profiles (MI300X)
This dataset contains PyTorch/Perfetto traces, offline execution logs, serving benchmarks, and comparative reports for Wan2.2-T2V-A14B sequence-parallel runs on AMD ROCm (gfx942, 8x MI300X node).
Dataset Directory Structure
reports/ / Root:
wan22_rocm_sp_sweep_analysis.md: 3-way sequence parallel topology comparison report.
wan22_profile_u4_r1_analysis.md: Detailed analysis of the Ulysses-4 topology.… See the full description on the dataset page: https://huggingface.co/datasets/Akshat/wan2.2-rocm-profiles.profile-assetsrobot-tts-profiles
Robot TTS Profiles (zero-train)
Damaged prompt clips + Chatterbox samples generated with cfg_weight=0.1 so the
model largely inherits the prompt's degradation instead of repairing it.
NOTE: cfg_weight=0.0 crashes chatterbox-tts 0.1.7 (t3.py hardcodes a CFG batch
of two); use 0.05-0.1 as the low-CFG setting.
prompts/ - one damaged reference clip per profile (scifi_robot, retro_8bit, intercom, telephone)
samples/ - generated speech per profile
base_clean.wav - the clean control… See the full description on the dataset page: https://huggingface.co/datasets/ACloudCenter/robot-tts-profiles.tweeter-profile
Dataset Card for tweeter-profile
** The original COCO dataset is stored at dataset.tar.gz**
Dataset Summary
tweeter-profile
Supported Tasks and Leaderboards
object-detection: The dataset can be used to train a model for Object Detection.
Languages
English
Dataset Structure
Data Instances
A data point comprises an image and its object annotations.
{
'image_id': 15,
'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB… See the full description on the dataset page: https://huggingface.co/datasets/Francesco/tweeter-profile.gemini-2.5-pro-tts-voice-profiles
Gemini 2.5 Pro TTS Voice Profiles
28,946 high-quality voice acting samples generated with Gemini 2.5 Pro Preview TTS, organized into 21 voice identities. Each sample is annotated with 59 Empathic Insight Voice Plus emotion/quality scores, BUD-E Whisper audio captions, and word-level timestamps. Includes per-voice FAISS similarity indices with GTE sentence embeddings for semantic search.
Repository Contents
File
Size
Description
{Voice}.tar (×21)
~1-2 GB… See the full description on the dataset page: https://huggingface.co/datasets/laion/gemini-2.5-pro-tts-voice-profiles.2D_profile
owner: Safran
license: cc-by-sa-4.0
data_production:
type: simulation
physics: 2D stationary RANS
simulator: elsA
num_samples:
test: 100
train: 300
storage_backend: hf_datasets
This dataset was generated with plaid, we refer to this documentation for additional details on how to extract data from plaid_sample objects.
The simplest way to use this dataset is to first download it:
from plaid.storage import download_from_hub
repo_id = "channel/dataset"
local_folder =… See the full description on the dataset page: https://huggingface.co/datasets/PhysArena/2D_profile.
