datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
uncgpt-training-datastallion-backup-uncgpt-20260713
UncGPT
A ~61 million parameter MoE-BitNet language model, pre-trained from scratch, designed to run on an ESP32-S3 Plus with 16 MB PSRAM at 2–4 tokens per second. The persona is a warm, grounded uncle — someone who sits down with you, calls the plumber on your behalf, orders groceries for a diabetic parent, coordinates a home-health aide, and stays with you when the day is heavy. No scripts. No hotlines. No dispatcher-speak. A person in the room, not a switchboard.
The target… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/stallion-backup-uncgpt-20260713.uncgpt-conversations-v7-69-total
UncGPT — Live-Approved Conversations (v7, 69-total)
4,761 approved multi-turn caregiving conversations across 11 languages, 69 skill axes, and 3 care levels. This is the live-approved canonical cohort used as the substrate for the NeurIPS 2026 UncGPT competition.
Part of the UncGPT NeurIPS 2026 Competition collection.
At a glance
Conversations
4,761
Languages
11 (en, yo, fr, pt, sw, zh, es, tl, bn, fa, hi)
Skill axes
69 (all covered)
Care levels… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-v7-69-total.uncgpt-conversations-semantic-approved-1p50-candidate
UncGPT — Semantic-Approved 1.50σ Conversations (Candidate)
The wider-tolerance (1.50σ) cohort against the same contrast semantic boundary. Useful as a higher-recall candidate for ablating gate strictness vs. coverage.
Part of the UncGPT NeurIPS 2026 Competition collection.
Configs
Config
What it is
approved_manifest (default)
conversations that passed at 1.50σ
rejected_manifest
conversations that failed even at 1.50σ
Why a wider tolerance
Some… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-semantic-approved-1p50-candidate.uncgpt-personas
UncGPT — Persona Pool
1,000,069 synthetic personas plus a 15,000-name registry (zero gaps in the build report). The substrate for all UncGPT caregiving conversation generation.
Part of the UncGPT NeurIPS 2026 Competition collection.
Configs
Config
Rows
What it is
preview (default)
10,000
first 10k personas — for the HF dataset viewer to render quickly
personas_full
1,000,069
complete v3 persona pool (3.5 GB; too large for the in-browser viewer)
names
15,000… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-personas.uncgpt-training-snapshot-20260504_phase_auncgpt-conversations-informal-approved-1p25
UncGPT — Informal-Register Approved
Conversations that passed the 1.25σ semantic gate AND the current strict programmatic gates — including intimate-register (tú-not-usted, tu-not-shoma, 你-not-您, no po/opo, plain not keigo), stricter colloquial Persian, and strict completion-integrity.
Part of the UncGPT NeurIPS 2026 Competition collection.
Counts
approved: 753
rejected: 1,475
skills covered: 53 of 69
by care: warm 450 / mid 152 / cold 151
by language: en 310 / sw… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-informal-approved-1p25.uncgpt-conversations-semantic-approved-1p25-repaired-paperclip
UncGPT — Semantic-Approved 1.25σ Conversations (Leak-Repaired)
The 1.25σ semantic-gate cohort with uncle-diary leakage repaired and normalized diary fields. The auditable replacement for the older paperclip_all_1803 source that an earlier audit flagged for visible diary leakage.
Part of the UncGPT NeurIPS 2026 Competition collection.
Config
approved_manifest (default): one row per approved conversation, with metadata + path back to the full-schema JSON.… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-semantic-approved-1p25-repaired-paperclip.uncgpt-persian-ultraclean-review-2026-05-14
UncGPT Persian Ultra-clean Review Subset
Viewer-friendly Persian / Arabic-script subset extracted from the current ultra-clean UncGPT training candidate corpus for manual inspection before training.
This dataset is for review/QC. It includes conversations selected when either:
the language hygiene audit marked the conversation as fa, or
Arabic-script characters appeared anywhere in seeker/Uncle text.
Splits / configs
conversations: one row per conversation, with… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-persian-ultraclean-review-2026-05-14.uncgpt-conversations-semantic-approved-1p25
UncGPT — Semantic-Approved 1.25σ Conversations
Conversations from the UncGPT v7 cohort that passed the tight (1.25σ) semantic gate against the contrast boundary. Useful for tight cohort training and as an ablation against the wider 1.50σ candidate cohort.
Part of the UncGPT NeurIPS 2026 Competition collection.
Configs
Config
What it is
approved_manifest (default)
rows for conversations that passed the 1.25σ semantic gate
rejected_manifest
rows for… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-semantic-approved-1p25.
