datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
armnet-demo-leaderboardmachine-failure-mlops-demo-logspiperx-demo558-value1500-a50-top10-union20-v1
PiperX advantage-selected teleoperation segments
Only pure human demonstrations. Value checkpoint step1500 (mixed demo+HIL); no HIL frames in this export.
A50 ranked globally across 558 source episodes, top10% AND A>0. Every selected start expands to [t,t+20); overlaps and adjacency merge. Each disconnected component is a separate output episode. Interior frames need not themselves be top10%.
Output: 7146 segments, 322261 frames, 2.983898 hours at30FPS.
Three camera streams and… See the full description on the dataset page: https://huggingface.co/datasets/Elvinky/piperx-demo558-value1500-a50-top10-union20-v1.UGround-Offline-Evaluationatomic-metrics-demographic-training-size
Atomic Metrics: Demographic Training-Size Analysis
Complete offline reproduction bundle for the effect of batch-selected training size on demographic preference prediction.
Version 2 — replaces the fixed-bank analysis. Select k extraction batches (five pairs each), use only their metrics and their 5k training pairs to refit BT/LR, then evaluate on cached test200 scores restricted to those metrics. Both the training rows and metric columns change with size. Extraction/refinement… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-demographic-training-size.commoncrawl-jobs-demo
Common Crawl on Jobs — datatrove JobsPipelineExecutor demo
228,088 English web documents (~820 MB compressed, 60 jsonl.gz shards) extracted from 1,237,374 Common Crawl pages — the output of a test run of datatrove's experimental JobsPipelineExecutor, which fans a datatrove pipeline out across a pool of Hugging Face Jobs instead of a Slurm cluster.
This is a pipeline demo artifact, not a curated corpus: one segment slice of one crawl, shared as the verifiable receipt for the… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/commoncrawl-jobs-demo.demoverse-personas-es-v1
Dataset Card for DemoVerse Personas ES v1
Resumen del dataset
demoverse-personas-es-v1 es un dataset de 100.000 personas sinteticas en espanol para Espana, disenado como artefacto publico y como capa operativa para simulacion sociológica.
El dataset se inspira metodologicamente en nvidia/Nemotron-Personas-France, pero no reutiliza sus filas ni intenta replicar la poblacion francesa. La adaptacion reescribe el marco para Espana, con clivajes territoriales, sistema de… See the full description on the dataset page: https://huggingface.co/datasets/apol/demoverse-personas-es-v1.trace-demolibero-rldx-demospeedup-slow2-fast4
LIBERO DemoSpeedup: RLDX-1 entropy, slow2 / fast4
Derived from kimtaey/libero_gr00t_delta using prehj/RLDX-1-IMG-LIBERO-60k.
Read FORMAT.md before training. Targets are explicit per-observation action chunks; the original action column is retained for provenance and is not the speedup target.
1693 episodes; 273465 source observations; 271772 usable anchors.
No model was trained for this release. See meta/demospeedup.json and per-episode entropy files.
reddit-demoReddit Demo dataset
reddit-demo
Reddit Demo dataset
Terminal_trajactory_demo
Terminal Agent Trajectory Demo
Complete multi-turn conversation trajectories of an AI agent solving programming tasks in a Linux terminal environment. Designed for training and evaluating Terminal/CLI agents.
Overview
Item
Details
Samples
20 (ID 1441–1460)
Task Language
Chinese instructions + English code
Difficulty
Medium
Expert Time Estimate
15 min
Environment
Linux / Python 3.13 / Docker
Task Categories
Category
Sample IDs… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/Terminal_trajactory_demo.2026-06-04-stwebagentbench-suitecrm-demos
2026-06-04-stwebagentbench-suitecrm-demos
Standing demo pool for Adversarial Inverse Constraint RL (ICRL) for LLM
orchestrator safety on ST-WebAgentBench (SuiteCRM easy tier). Every
experiment run consumes this pool; per-run artifacts (embeddings, constraint
heads, adapters, CuP evals) live in separate <date>-<run-name> repos in this
namespace.
field
value
experiment
ICRL safe/unsafe demo pool: constraint C_theta is learned from the safe demos only; unsafe demos are… See the full description on the dataset page: https://huggingface.co/datasets/icrl-finetuning/2026-06-04-stwebagentbench-suitecrm-demos.every-eval-ever-demopropaganda_demonizace
Information
This is a dataset extracted from MU-NLPC/Propaganda dataset. Please refer to original dataset for further information (including licensing).
Meaning of Attributes
The definition of each propaganda technique is labeled in annotation manual here. More information can be found at the dataset website.
arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard index: 1
Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416.arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard index: 1
Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416.reddit-demoarxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard index: 1
Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417.synthetic_demographics_seed
Synthetic Demographic Seeds v1
This is a dataset of 3,541,040 roughly demographically correct demographic seeds and somewhat demographically accurate names all generated from publicly available datasets.
(note there were tradeoffs made with accuracy and what I could tie together, v2 will be more accurate)
get_synthetic_demographics.py contains a method for quickly and randomly selecting batches of demographic seeds.
There is no filtering on this at the moment.
Format… See the full description on the dataset page: https://huggingface.co/datasets/sacrificialpancakes/synthetic_demographics_seed.massive_serve_demodemo-experiment-two
Demo Experiment
A demo experiment to test the builder flow with various question types.
Dataset Overview
Property
Value
Run ID
036df3a6-6f8c-45fb-979d-be4f990bd0bf
Status
completed
Created
12/23/2025, 9:11:08 PM
Generator
LocalBench v0.1.0
Statistics
Metric
Value
Total Generations
10
Successful
10 (100.0%)
Failed
0
Average Latency
2803ms
Total Duration
25.2s
Configuration
Models… See the full description on the dataset page: https://huggingface.co/datasets/GhostScientist/demo-experiment-two.arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416.demo-datasets
Nixiesearch demo datasets
This repo has a couple of small-scale demo datasets in JSON format you can use for playing with nixiesearh:
MSRD: A TMDB-sourced dataset of popular movies.
MSRD: A movie search ranking dataset
See the original dataset docs for more details
Example record:
{
"_id": "27205",
"title": "Inception",
"overview": "Cobb, a skilled thief who commits corporate espionage by infiltrating the subconscious of his targets is offered a chance to regain… See the full description on the dataset page: https://huggingface.co/datasets/nixiesearch/demo-datasets.toca_demo_dataset
TOCA Demo Dataset (LeRobot Compatible)
This dataset was converted from the BOTA driver SQLite recording.
Structure
observation.state: Joint positions (6 DOF)
observation.velocity: Joint velocities (6 DOF)
observation.wrench: Force/Torque sensor data (Fx, Fy, Fz, Tx, Ty, Tz)
action: Joint positions (Target for Imitation Learning)
episode_index: ID of the episode
frame_index: Frame number within the episode
timestamp: Time in seconds
next.done: Boolean indicating end of… See the full description on the dataset page: https://huggingface.co/datasets/gribes02/toca_demo_dataset.thirawat-mapper-demo-indexarxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard index: 1… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416.Tweet_demoarxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-20260416.gambit-demo-projecttasks
Gambit Demo Project Tasks
A small demo dataset of project task records used to validate the Gambit AI Platform ingestion pipeline and retrieval scenarios. Intended for getting-started tutorials, smoke tests, and reproducible examples.
Disclaimer
Records originate from an internal project plan template (PMO_NAV1, "Proje Yönetim Ofisi NAV Proje Planı Şablonu") used as a generic blueprint.
The dataset contains no real customer data. customerName is null in every record.
All… See the full description on the dataset page: https://huggingface.co/datasets/ylmzunl/gambit-demo-projecttasks.
