datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SLT-Task2-Post-ASR-Speaker-Tagging
Dataset Name: Dataset for ASR Speaker-Tagging Corrections (Speaker Diarization)
Description
This dataset is pairs of erroneous ASR output and speaker tagging, which are generated from a ASR system and speaker diarization system.
Each source erroneous transcription is paired with human-annotated transcription, which has correct transcription and speaker tagging.
SEGment-wise Long-form Speech Transcription annotation (SegLST), the file format used in the CHiME challenges… See the full description on the dataset page: https://huggingface.co/datasets/GenSEC-LLM/SLT-Task2-Post-ASR-Speaker-Tagging.postgresql-llm
postgresql-llm
A pure PostgreSQL dataset for training and evaluating LLMs on PostgreSQL SQL and PL/pgSQL. Every row is a (question, schema, SQL) triplet with rich metadata for filtering and analysis.
Dataset Summary
postgresql-llm is a pure PostgreSQL dataset: SQL and PL/pgSQL only, with metadata for difficulty, category, and source.
Metric
Value
Total rows
211,539
PostgreSQL-specific rows
11,998 (5.7%)
Schema fill rate
82.2%
Explanation fill rate
17.8%… See the full description on the dataset page: https://huggingface.co/datasets/neurondb/postgresql-llm.mu-glioma-post-processed
Processed MU-Glioma-Post
Start Here
Use manifests/experiment_index.csv as the main case-level table.
Use manifests/longitudinal_index.csv when ordering patient timepoints for longitudinal work.
Use manifests/validation_summary.json to confirm the processed tree is complete.
Directory Guide
images_native/
Canonical modality symlinks to the source images.
File names use t1, t1c, flair, and t2.
images_reoriented/
Not populated because audit showed all volumes… See the full description on the dataset page: https://huggingface.co/datasets/sbandred/mu-glioma-post-processed.postdyn-artifactsSAND-Post-Training-Dataset
SAND-Post-Training-Dataset: High-Quality Synthetic Reasoning Dataset Built with AMD GPUs
Dataset Summary
We introduce the SAND-Post-Training-Dataset, a high-quality synthetic reasoning dataset for mathematics and science built entirely using a synthetic data pipeline running on the AMD ROCm™ stack and AMD Instinct™ MI325 GPUs.
This dataset prioritizes difficulty and novelty over volume, demonstrating that high-difficulty synthetic data can elevate… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-Post-Training-Dataset.turkish_political_position_benchmark
Turkish Political Position Benchmark
The Turkish Political Position Benchmark measures how language models respond to normative statements about Turkish politics. It reports ideological dimension scores and response similarity to documented political-party reference profiles.
The benchmark does not claim that a model belongs to a party, has a voting intention, or possesses political beliefs. A party similarity score only means that the model produced a similar pattern of answers… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/turkish_political_position_benchmark.chess-rl-evalbackln-guest-post-quality-public-mirror
Backln Guest Post Quality Public Mirror
Public-safe mirror for validating Hugging Face Dataset Viewer indexing and release gates. This dataset is not the private training corpus.
Full text, titles, and snippets are removed by default. The mirror keeps labels, coarse metadata, feature buckets, and hash prefixes so the public Hub can verify schema and distribution without exposing customer content.
Schema
label: one of published, manual_review, rejected.
source: coarse… See the full description on the dataset page: https://huggingface.co/datasets/driodnexus/backln-guest-post-quality-public-mirror.mpii-human-pose-captions
Dataset Card for MPII Human Pose Descriptions
Dataset Summary
The MPII Human Pose Descriptions dataset extends the widely-used MPII Human Pose Dataset with rich textual annotations. These annotations are generated by various state-of-the-art language models (LLMs) and include detailed descriptions of the activities being performed, the count of people present, and their specific poses.
The dataset consists of the same image splits as provided in MMPose, with 14644… See the full description on the dataset page: https://huggingface.co/datasets/saifkhichi96/mpii-human-pose-captions.postmortems-exploitsstego-bench
TODO (maintainer): confirm the final license (the license: field above is a placeholder) and update the LICENSE file before flipping the gate to public. The rest of this card assumes the repo stays gated: manual.
StegoBench
Evaluating steganography potential in language models through supervised learning.
This repository hosts the artifacts that accompany the paper StegoBench: Evaluating steganography potential in language models through supervised learning (NeurIPS 2026… See the full description on the dataset page: https://huggingface.co/datasets/Poseidon-Research/stego-bench.2026-08-26-sonnet45-post-action-retrospection-natural-turn-design
synth post_action_retrospection run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth post_action_retrospection run — per-stage snapshots (resumable generation cache)
date_generated
20260826_152715
constitution
constitutions/claude_distilled_12_principles_mid/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ c2fdee460e71fa28e9902edf1cc662db0d19cad8
models
per-stage models — see… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-26-sonnet45-post-action-retrospection-natural-turn-design.PostureAndSustainmentOptimization
Posture And Sustainment Optimization
This dataset contains research artifacts for Posture and Sustainment Optimization, a benchmark and simulation project for optimizing distributed posture, readiness, and sustainment decisions under uncertainty.
Source repository: https://github.com/anote-ai/research-postureandsustainmentoptimization
Displayable Configs
The Hugging Face viewer reads normalized JSONL tables under viewer/:
exp1_metrics: greedy placement baseline… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/PostureAndSustainmentOptimization.OpenO1_SFT_ultra_BoN_positvie_reward_v3_N-sampleksl-pose-dictionary-poc
KSL Pose Dictionary (PoC)
한국수어(KSL) text-to-pose 시제품용 keypoint 데이터셋.
docent_AI_sign_research_02 프로젝트에서 생성. Neural Sign Actors (CVPR 2024) 접근법을 KSL에 적용하는 Path B (Dictionary-based) 시제품의 핵심 데이터셋.
개요
자산
갯수
키포인트
sldict keypoint (국립국어원 한국수어사전)
1,444 단어
OpenPose 137 (RTMW-DW-L-M 추출)
NIASL2021 gloss segmentation keypoint (재난 안전 도메인)
2,287 base gloss
OpenPose 137 (NIASL 원본)
Hybrid sign index
4,511 unique signs
단어 → keypoint 경로 매핑
Stage 1 학습 corpus
20,085 samples… See the full description on the dataset page: https://huggingface.co/datasets/Trotquonalize/ksl-pose-dictionary-poc.PosTWITAdqs-post-training
DQS Post-Training Preference Data
Strict English-to-Korean preference data for three post-training objectives.
All three configurations contain the same ordered set of 5,200 preference
examples after source-quality review and exclusion of one Teacher/Student pair
with no response-level preference.
Run-prefixed layout
The original root-level mpo/, cpo/, dpo/, and manifest.json are retained as the legacy Gemma release for compatibility with existing download… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/dqs-post-training.PosterBench
PosterBench
PosterBench is a 100-paper, image-native benchmark for academic poster
generation introduced in AutoDesign: Meta-Harness Optimization for
Long-Horizon Agentic Design.
This Hugging Face release is metadata only. It does not host or redistribute
the underlying paper PDFs, full text, abstracts, figures, tables, or
thumbnails. Each row identifies the exact benchmark paper version and records
the official landing page, access policy, license information where available… See the full description on the dataset page: https://huggingface.co/datasets/YaxinLuo/PosterBench.post-cutoff-2024-2026-bundles
post-cutoff-2024-2026-bundles
12 research briefings (53,685 words / ~70K tokens) covering events from April 2024 through May 2026. Built as source material for context-distillation SFT of a pre-April-2024 base model, and usable directly as a small CPT-style corpus.
Format
{
"text": "<full markdown bundle>",
"topic": "ai_ml_2024_2026",
"word_count": 4950,
"char_count": 37474
}
Each bundle is markdown with ###-level entries (typically 10–18 entries per bundle)… See the full description on the dataset page: https://huggingface.co/datasets/Lambent/post-cutoff-2024-2026-bundles.repro-siamesenorm-breaking-the-barrier-to-reconciling-pre-post-norm-traces
Agent traces
Agent sessions published from a Trackio Logbook.
PosterBench-mini
PosterBench-mini
PosterBench-mini is the fixed 10-paper development subset of PosterBench,
introduced in AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic
Design. It contains exactly two papers from
each of the benchmark's five disciplines.
This Hugging Face release is metadata only. It does not host or redistribute
the underlying paper PDFs, full text, abstracts, figures, tables, or
thumbnails.
Load
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/YaxinLuo/PosterBench-mini.hf-posts
Hugging Face Posts
This dataset contains posts scraped from https://huggingface.co/posts.
It includes all posts published from the launch date on December 23, 2023, up to November 24, 2024, at 15:40.
sentiment_postsPretergeek__OpenChat-3.5-0106_32K-PoSE-details
Dataset Card for Evaluation run of Pretergeek/OpenChat-3.5-0106_32K-PoSE
Dataset automatically created during the evaluation run of model Pretergeek/OpenChat-3.5-0106_32K-PoSE
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pretergeek__OpenChat-3.5-0106_32K-PoSE-details.chess-decoy-positions20th-century-postage-stamp
This 20th-century postage stamp collection once belonged to Tomás Rudass, a Colombian sailor who traveled the world and collected memories from his journeys.
He acquired stamps from the countries he visited as a way of preserving the places, cultures, and experiences he encountered along the way. As a historian,
I have undertaken the task of digitizing this collection in order to preserve his memory and make it accessible for research, education, and public exploration.
This… See the full description on the dataset page: https://huggingface.co/datasets/Vault-of-History/20th-century-postage-stamp.hackaday-posts
🚀 Hackaday Universe: 50K+ Tech Articles & Vibrant Maker Conversations
Dive into the ultimate collection of Hackaday's tech universe! This isn't just another dataset—it's a living archive of maker culture, featuring 54,599+ articles with complete comment threads where brilliant minds collide, debate, and innovate together.
🔥 Why This Dataset Rocks
🤖 Perfect for AI Training
Train models on authentic technical writing and community interactions
Learn from real… See the full description on the dataset page: https://huggingface.co/datasets/hybridfree/hackaday-posts.QualityVision-Locomotion-Pose-Dataset-Walking-Jogging-Running
QualityVision Locomotion Pose Dataset (Walking + Jogging + Running) — Sample
This is a compact, viewer-friendly sample extracted from a much larger HQ locomotion export generated by the QualityVision Motion Dataset Engine.
Looking for the full commercial export or custom delivery? See pricing & ready-made bundles on qvision.space/dataset-pricing.
What’s inside
data.jsonl: one JSON object per line (one frame per row) with 33 MediaPipe/BlazePose landmarks (x,y,z… See the full description on the dataset page: https://huggingface.co/datasets/Alaaharoun/QualityVision-Locomotion-Pose-Dataset-Walking-Jogging-Running.aws-rag-qa-positives
QA Pairs from AWS Service Documentation
This dataset contains chunked documentation AWS service documentation, along with QA pairs generated from the chunked content. The following AWS services have excerpts of documentation in this dataset:
AWS Lambda
AWS RDS
AWS EC2
AWS ECS
AWS ECS Fargate
AWS Elastic Beanstalk
AWS EKS
AWS Wavelength
AWS Outposts
AWS Bedrock
AWS Sagemaker
AWS QBusiness
AWS QDeveloper
AWS Batch
AWS API-Gateway
AWS Cloudfront
AWS Athena
AWS Aurora
AWS Dynamo DB
AWS… See the full description on the dataset page: https://huggingface.co/datasets/CadenShokat/aws-rag-qa-positives.QualityVision-Jogging-Pose-Dataset-61-Videos-14550-Frames
QualityVision Jogging Pose Dataset (61 videos, 14,550 frames) — Sample
This Hugging Face dataset is a compact sample extracted from the full QualityVision Jogging Pose export.
Action label: jogging
Keypoints: 33 landmarks per person (MediaPipe / BlazePose) with x, y, z, visibility
Post-processing (as exported): temporal smoothing + body normalization flags are included in metadata
Use this sample to validate the schema and quality before purchasing larger exports.
Pricing &… See the full description on the dataset page: https://huggingface.co/datasets/Alaaharoun/QualityVision-Jogging-Pose-Dataset-61-Videos-14550-Frames.
