datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
state-cancer-profiles
United States State Cancer Profiles data extract (mirror)
This is a mirror. Cite the Zenodo record, not this page:
Davis S. United States State Cancer Profiles data extract — vintage V3. Zenodo. https://doi.org/10.5281/zenodo.22085273
Concept DOI (always resolves to the latest vintage): https://doi.org/10.5281/zenodo.11098814
No DOI is minted on Hugging Face. HF hosts these bytes for native hf:// / DuckDB access and an ML audience that would never find the Zenodo record;… See the full description on the dataset page: https://huggingface.co/datasets/seandavis/state-cancer-profiles.laion-voice-profiles-annotated
LAION Voice Profiles — Annotated
Authors: Christoph Schuhmann and LAION.
28,212,933 utterances / 71,056 hours of synthetic English and German voice-acting speech from
500 distinct voice profiles, each driven through the same fixed matrix of 842 named acting
conditions. Every
utterance carries 40 emotion intensities, 57 perceptual voice dimensions, 4 audio-quality heads,
vocal-burst detections with timings, word-level forced alignment, MOSS audio codec tokens, a
768-d… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-annotated.voice-profileslaion-voice-profiles-dpo
LAION Voice Profiles — TTS preference pairs (DPO)
Authors: Christoph Schuhmann and LAION.
3,451,531 preference pairs in three families, built from the same 500 synthetic voice profiles
as laion/laion-voice-profiles-sft.
Every pair shares one prompt; chosen and rejected differ only in the way the family names.
config
pairs
teaches
how rejected is made
emotion
1,064,594
hit the right tone for this line
a take of the same sentence, same voice, rendered at the wrong… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-dpo.laion-voice-profiles-dpo-cfg
LAION Voice Profiles — contrastive DPO pairs (CFG + phase 2)
Authors: Christoph Schuhmann and LAION.
842,935 preference pairs in four families, built from the same 500 synthetic voice profiles as
laion/laion-voice-profiles-sft
and laion/laion-voice-profiles-dpo.
These are the two pair families that the sister DPO set does not contain: they were built later,
for two measured defects of the models trained on it, and they are the complete remainder of the
project's preference… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-dpo-cfg.profileslaion-voice-profiles-sft
LAION Voice Profiles — TTS supervised fine-tuning set
Authors: Christoph Schuhmann and LAION.
1,200,531 instruction-tuning samples for reference-conditioned TTS, drawn from 500 synthetic
voice profiles: for each of the 842 acting conditions of each voice, the best 3 of its 48
candidate takes, each paired with a different clip of the same voice from another group as
the reference, plus the exact conditioning prompt, MOSS audio codes for both target and reference,
word-level… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-sft.recruitment-dataset-candidate-profiles-english
Djinni Dataset (English CVs part)
Overview
The Djinni Recruitment Dataset (English CVs part) contains 150,000 job descriptions and 230,000 anonymized candidate CVs, posted between 2020-2023 on the Djinni IT job platform. The dataset includes samples in English and Ukrainian.
The dataset contains various attributes related to candidate CVs, including position titles, candidate information, candidate highlights, job search preferences, job profile types, English… See the full description on the dataset page: https://huggingface.co/datasets/lang-uk/recruitment-dataset-candidate-profiles-english.User_Profiles_MBTIThis dataset include 85462 Users profiles with MBTI personality traits from Personality Cafe, , with the following information:
Usernames
MBTI types
Gender
Followers
Self-descriptions (About section)
Sexual orientation
Enneagram Type
There is a small version of this dataset of 17,000 users with integrated personalties, you can find it at the Github
🌹Please Cite Our Work If Helpful:
Thanks! / 谢谢! / ありがとう! / merci! / 감사! / Danke! / спасибо! / gracias! ...
@inproceedings{shu2024llm… See the full description on the dataset page: https://huggingface.co/datasets/ZoeyShu/User_Profiles_MBTI.gemini-2.5-pro-tts-voice-profiles-prompts
Gemini 2.5 Pro TTS Voice Profiles — with performance prompts
This is laion/gemini-2.5-pro-tts-voice-profiles
augmented with two voice-acting performance prompts per sample, generated by
google/gemma-3-12b-it from the dataset's existing annotations (BUD-E Whisper
caption, Empathic-Insight scores, Gemini word-level transcriptions/captions, and
vocal-burst descriptions) — no re-running of ASR or whisper experts.
Everything from the original dataset is preserved; each sample's .json… See the full description on the dataset page: https://huggingface.co/datasets/laion/gemini-2.5-pro-tts-voice-profiles-prompts.qwen36-coding-layer-bottleneck-profiles
Qwen3.6 Coding Layer Bottleneck Profiles
This dataset packages a local llama.cpp/ATX profiling campaign for identifying which whole transformer layers are the strongest candidates to keep hot for coding and agentic workloads.
The goal is to compare a learned top-layer policy against architecture heuristics such as the actual full-attention layers, first-10, and last-10. The included results are timing-attribution measurements, not CUDA speedup claims. They are intended to seed… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/qwen36-coding-layer-bottleneck-profiles.mazinger-dubber-profiles
Mazinger Dubber — Voice Profiles
Voice profiles for mazinger-dubber. Hosted on HuggingFace:
https://huggingface.co/datasets/bakrianoo/mazinger-dubber-profiles
Adding a New Profile
1. Prepare your files
Create a folder named after the profile:
profiles/
└── my-name/
├── script.txt # Plain-text transcript matching the audio exactly
└── voice.m4a # Voice sample (supported: .m4a, .wav, .mp3)
Tips: 10–30 seconds of clear speech, minimal… See the full description on the dataset page: https://huggingface.co/datasets/bakrianoo/mazinger-dubber-profiles.ray-data-gpu-idle-profiles
Ray Data GPU Idle Profiles (B200)
Nsight Systems profiles (exported to SQLite, readable by
nsys-ai) from an experiment on how a
Ray Data pipeline keeps a GPU idle, and how the loss splits between moving
data and waiting for data. Captured on a single NVIDIA B200 with Ray 2.58.0
/ master, PyTorch 2.14.0+cu130, Nsight Systems 2026.1.3.
These profiles back the write-up in the iThome Ironman series
「GPU 很忙?他真的有在做事嗎?」 (Days 27–29), and are shared so the numbers and
the before/after… See the full description on the dataset page: https://huggingface.co/datasets/rich7421/ray-data-gpu-idle-profiles.gemini-2.5-pro-tts-voice-profiles
Gemini 2.5 Pro TTS Voice Profiles
28,946 high-quality voice acting samples generated with Gemini 2.5 Pro Preview TTS, organized into 21 voice identities. Each sample is annotated with 59 Empathic Insight Voice Plus emotion/quality scores, BUD-E Whisper audio captions, and word-level timestamps. Includes per-voice FAISS similarity indices with GTE sentence embeddings for semantic search.
Repository Contents
File
Size
Description
{Voice}.tar (×21)
~1-2 GB… See the full description on the dataset page: https://huggingface.co/datasets/laion/gemini-2.5-pro-tts-voice-profiles.wan2.2-rocm-profiles
Wan2.2 Sequence Parallel ROCm Profiles (MI300X)
This dataset contains PyTorch/Perfetto traces, offline execution logs, serving benchmarks, and comparative reports for Wan2.2-T2V-A14B sequence-parallel runs on AMD ROCm (gfx942, 8x MI300X node).
Dataset Directory Structure
reports/ / Root:
wan22_rocm_sp_sweep_analysis.md: 3-way sequence parallel topology comparison report.
wan22_profile_u4_r1_analysis.md: Detailed analysis of the Ulysses-4 topology.… See the full description on the dataset page: https://huggingface.co/datasets/Akshat/wan2.2-rocm-profiles.pythia-deduped-memorisation-profilesThis dataset has been created as an artefact of the paper Causal Estimation of Memorisation Profiles (Lesci et al., 2024).
More info about this dataset in the related collection Memorisation-Profiles.
personalaity-llm-personality-profiles
PersonalAIty: HEXACO personality profiles of frontier LLMs
Self-reported HEXACO personality profiles for 10 frontier language models across 8 vendors,
measured on 2026-08-16 with an open 50-item inventory, plus the instrument itself so the
measurement can be rerun or criticised.
This is a snapshot with a date on it, not a standing benchmark. Model versions drift; the
value here is that the whole measurement is reproducible with one command against models anyone
can reach.… See the full description on the dataset page: https://huggingface.co/datasets/Sciupy/personalaity-llm-personality-profiles.recruitment-dataset-candidate-profiles-ukrainian
Djinni Dataset (Ukrainian CVs part)
Overview
The Djinni Recruitment Dataset (Ukrainian CVs part) contains 150,000 job descriptions and 230,000 anonymized candidate CVs, posted between 2020-2023 on the Djinni IT job platform. The dataset includes samples in English and Ukrainian.
The dataset contains various attributes related to candidate CVs, including position titles, candidate information, candidate highlights, job search preferences, job profile types, English… See the full description on the dataset page: https://huggingface.co/datasets/lang-uk/recruitment-dataset-candidate-profiles-ukrainian.people-profiles-io-salesOkCupid-59k-Anonymized-Profiles
💘 OkCupid 59k Anonymized Profiles
This dataset contains 59k anonymized OkCupid dating profiles, converted from the original CSV dataset into Parquet format.
It includes structured profile attributes such as age, gender, orientation, body type, lifestyle habits, education, job, location, and several free-text essay fields written by users.
✍️ Essay Fields
The columns essay0 to essay9 correspond to open-ended profile questions from OkCupid.
These fields contain natural… See the full description on the dataset page: https://huggingface.co/datasets/SpiceeChat/OkCupid-59k-Anonymized-Profiles.ami-long-form-profilesffhpt5921-1d1aab-customer-profiles
Customer Profiles
Customer demographic and contact profiles collected from opt-in sources.
Overview
Brief description of the dataset, its contents, and intended use.
Governance
This dataset is governed by the Data Platform compliance policy.
License: cc-by-4.0
Owner Team: marketing
Classification: restricted
Retention Days: 730
Contains PII: Yes
africa-synth-education-teacher-profiles-nigeria
Nigeria Education - Teacher Profiles | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: education - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Education… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-education-teacher-profiles-nigeria.ARGO_ProfilesArgovis Argo Ocean Profiles
Dataset summary:
This dataset contains ocean profile data collected by the international Argo float program and accessed via the Argovis API. Each record corresponds to a single profile measured by an autonomous drifting float, including time, location, basin, and associated profile metadata fields that can be joined to the underlying temperature and salinity data structures. The goal of this dataset is to provide a ready-to-use subset of Argo profiles… See the full description on the dataset page: https://huggingface.co/datasets/Otter21/ARGO_Profiles.pearl_with_profilesprescreen-profilesafrica-synth-retail-and-ecommerce-customer-profiles-nigeria
Customer Profiles | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-retail-and-ecommerce-customer-profiles-nigeria.cloudcanvas_ast001_customer_profiles
Customer Profiles
Quarterly snapshot of customer profile attributes used by the CloudCanvas data platform.
indic-synthetic-profiles
🇮🇳 Indian Synthetic Identity Dataset
10,000 realistic Indian synthetic identities across 8 languages — generated by indic-faker
Dataset Description
This dataset contains 10,000 rows of realistic, synthetic Indian identity data generated using the indic-faker Python library. Every record is algorithmically valid — Aadhaar numbers pass Verhoeff checksum verification, GSTINs have correct state codes, and names are culturally authentic across 8 Indian languages.… See the full description on the dataset page: https://huggingface.co/datasets/adwaith06/indic-synthetic-profiles.bluesky_profiles
Bluesky Network (Profiles and Follows)
This is a scraped mirror of the Bluesky (https://bsky.app/) social graph. It includes profile information (did, handle, display name, indexed at, follows count, followers count, posts count, and descriptions). The follow graph is (did, did) relationships, with created at timestamp. There is also a calculated PageRank of the follows graph.
Notes:
Consult the Bluesky / AT Proto API docs for explainations for fields.
Scraping prioritizes larger… See the full description on the dataset page: https://huggingface.co/datasets/andrewconner/bluesky_profiles.
