datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task1192_food_flavor_profile
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1192_food_flavor_profile
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1192_food_flavor_profile.dpo_profile
Multi-Personality Generation (MPG) Datasets
Paper | Code
This repository contains datasets released as part of the Multi-Personality Generation (MPG) framework. MPG is a decoding-time paradigm that enables Large Language Models (LLMs) to simultaneously embody multiple personalization attributes without requiring extra training or multi-dimensional models.
Dataset Description
The collection includes several Direct Preference Optimization (DPO) datasets used for MBTI… See the full description on the dataset page: https://huggingface.co/datasets/RongxinChen/dpo_profile.indic-synthetic-profiles
🇮🇳 Indian Synthetic Identity Dataset
10,000 realistic Indian synthetic identities across 8 languages — generated by indic-faker
Dataset Description
This dataset contains 10,000 rows of realistic, synthetic Indian identity data generated using the indic-faker Python library. Every record is algorithmically valid — Aadhaar numbers pass Verhoeff checksum verification, GSTINs have correct state codes, and names are culturally authentic across 8 Indian languages.… See the full description on the dataset page: https://huggingface.co/datasets/adwaith06/indic-synthetic-profiles.virtual-patient-profiles-sample
VHS Patient Profiles — Nutrition Education Simulation
Dataset Summary
A curated collection of 8 richly structured synthetic patient personas designed for healthcare education simulations, with a primary focus on dietetics and nutrition counseling training. Each record represents a complete clinical case with demographic, clinical, psychosocial, and behavioral dimensions, along with structured guidance for educators and AI simulators.
Profiles are intended for use… See the full description on the dataset page: https://huggingface.co/datasets/joeljames270/virtual-patient-profiles-sample.daniel-os-profile-sft
Daniel OS Profile SFT and Behavior Tests
Small, source-grounded datasets used to adapt and evaluate the browser-native
Daniel OS portfolio assistant. The model separates Daniel-specific claims from
general definitions, synthesizes definitions from retrieved evidence, requests
public retrieval when evidence is absent, and declines private-person requests.
Splits
Configuration
Split
Records
Purpose
sft
train
268
Profile-grounded conversational fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/danelcsb/daniel-os-profile-sft.kakoverse-crisis-profiles-v0
KakoVerse Crisis Profiles v0
Age-indexed crisis summaries (ages 20-100) aligned with the persona grid.
Usage
from datasets import load_dataset
ds = load_dataset("Reza2kn/kakoverse-crisis-profiles-v0")
print(ds["train"][0])
Provenance
Generator version: Scripts/publish_to_hf.py
Contact: https://huggingface.co/Reza2kn
deepmath-l5-9-qwen3-4b-at8-profile
DeepMath-L5-9 10K @8 Profiled by Qwen3-4B-Instruct-2507
Difficulty-stratified rollout profile of 10,000 DeepMath problems (levels 5-9)
by Qwen3-4B-Instruct-2507, 8 rollouts per question (n=8, T=0.7,
top_p=0.95, max_tokens=16384).
Built for the OPSD context-strength study: comparing two OPSD context
sources (gold ref-solution vs hint sequence) across three difficulty buckets.
The core hypothesis: when the prepended context is too strong, OPSD degrades
into SFT — the student is just… See the full description on the dataset page: https://huggingface.co/datasets/chichi56/deepmath-l5-9-qwen3-4b-at8-profile.profile-qa-synthetic-public-v1
Profile-QA Synthetic Public V1
Description
This dataset contains deterministic synthetic Q&A examples for public
resume/profile answering. It was generated from generic resume sections and
public-style facts, with evidence references back to section_id and fact_id.
The ontology is intentionally reusable across people and forks:
identity, current_role, experience, projects, education,
recommendations, skills, and interests. Temporal and practical sections
are… See the full description on the dataset page: https://huggingface.co/datasets/justinthelaw/profile-qa-synthetic-public-v1.persona-profiles-1m
