datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Sekai2_Real_World
Sekai2 Real World
This repository releases the reproducible URL/timestamp metadata and paired
camera-pose/caption annotations for the perspective-video portion of
Sekai2. See the paper: Sekai2: From World Exploration to Interactive World Modeling.
Resources: 🌐 Project Page · 💻 GitHub · 📄 Paper
The perspective MP4 clips are not redistributed here. Each row in
sekai2_clips.csv provides the source URL and the exact half-open frame range
[start_frame, end_frame) in a canonical 30… See the full description on the dataset page: https://huggingface.co/datasets/Kangverse/Sekai2_Real_World.privacy-preserving-real-world-human-motion-sample
Privacy-Preserving Real-World Human Motion Sample
A market-validation sample of anonymous 2D skeleton/pose observations derived from a real-world indoor CCTV stream.
Why this sample exists
We are validating demand for continuously collected, privacy-oriented real-world human-motion data before expanding to multi-camera releases.
Current public sample
750 public observations
derived pose/skeleton data
anonymous track identifiers
no raw RGB video
no… See the full description on the dataset page: https://huggingface.co/datasets/Ragab-Adel/privacy-preserving-real-world-human-motion-sample.innoduel-rlhf-real-world-human-preferences-sample
Real-World Human Pairwise Preferences — Public Sample
📦 This is a free, public sample of a commercial dataset.
It contains 1,350 rows curated for inspection. The full dataset has 1.5 million
human pairwise-preference decisions.
Full dataset: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf
Request access / licensing: see § Access to the full dataset — contact kari.nieminen@nordo.fi.
Use this sample to evaluate the data's quality, structure and… See the full description on the dataset page: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf-real-world-human-preferences-sample.real-world-benign-use-cases
Real-World Benign Use Cases
A curated set of 178 real-world, 100%-benign examples (label == 0 for every row) pulled from
production AI-coding-agent traffic — chat messages, tool output, shell commands, code snippets —
built specifically to stress-test prompt-injection / jailbreak classifiers for false positives.
Every row was independently judged benign with high confidence before inclusion. This is not a
random sample of production traffic: rows were preferentially drawn from… See the full description on the dataset page: https://huggingface.co/datasets/rogue-security/real-world-benign-use-cases.abhijitdahatonde_real-world-smartphones-dataset
Real World Smartphone's Dataset
Worlds Smartphones: A Comprehensive Dataset for Cutting-Edge Analysis
Dataset Info
Source: Kaggle
Original Size: 0.02 MB
Kaggle Downloads: 5,744
Files: 1
Files
smartphones.csv
Mirrored from Kaggle
clinical-quad-trial-pop-variance-realworld-subgroup-signal-generalization-claim-drift-v0.1What this repo does
This dataset models population mismatch narrative drift in clinical trial reporting. It predicts when the interaction between trial population variance, real-world variance, subgroup signal strength, and generalization claim intensity indicates that narrative claims extend beyond what the data supports.
Core quad
trial_population_variance_index
real_world_variance_index
subgroup_signal_strength_index
generalization_claim_index
Prediction target
label_claim_drift
Row… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-trial-pop-variance-realworld-subgroup-signal-generalization-claim-drift-v0.1.legal-law-real-world-coherence-decay-v0.1What this dataset is
You receive
rule text
real world conditions
enforcement pattern
evasion signals
reform pressure
You decide
Does the law still align with reality
Answer
coherent
or
incoherent
Why this matters
When law diverges from reality
enforcement weakens
reform pressure rises
doctrine shifts
This dataset measures that divergence early.
realworld-ai-support-dialog-benchmark-v1
Real-World AI Support Dialog Benchmark v1
This dataset is a synthetic but realistic benchmark for evaluating AI assistants in support workflows.
Why this dataset exists
Many AI demos are too toy-like to reflect production support conversations. This benchmark simulates realistic support cases with:
Ambiguous user intent
Multi-turn clarifications
Policy constraints (refund windows, account security, compliance)
Escalation and handoff conditions
Hallucination-risk prompts… See the full description on the dataset page: https://huggingface.co/datasets/egroupai/realworld-ai-support-dialog-benchmark-v1.amazon_realworld_250_rows_dataset
