datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xlam_hermes_validated5_synt_flux_street_selected_single_validated_1011synthvision-validated-qwen-by-kimi
synthvision-validated-qwen-by-kimi
Qwen 3.5 annotations validated by Kimi K2.5 (93.1% pass rate)
Records: 55,359
About
Cross-validated subset from the SynthVision pipeline. Kimi K2.5 reviewed all 59,476 Qwen 3.5 annotations and confirmed 55,359 as consistent with the source images (93.1% pass rate).
Validation criteria: consistent == true AND confidence >= 0.7. Records that failed validation were removed — primarily cases where the annotator hallucinated findings not… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/synthvision-validated-qwen-by-kimi.5_synt_flux_street_selected_single_validated_v2_1011synthvision-validated-kimi-by-qwen
synthvision-validated-kimi-by-qwen
Kimi K2.5 annotations validated by Qwen 3.5 (93.0% pass rate)
Records: 55,382
About
Cross-validated subset from the SynthVision pipeline. Qwen 3.5 reviewed all 59,539 Kimi K2.5 annotations and confirmed 55,382 as consistent with the source images (93.0% pass rate).
Validation criteria: consistent == true AND confidence >= 0.7. Records that failed validation were removed — primarily cases where the annotator hallucinated findings not… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/synthvision-validated-kimi-by-qwen.prima-gold100-validated5_synt_flux_landmark_missing_imgs_validatedSWE-bench_validated_12_18_style-3__fs-oracleSWE-bench_validated_12_18_bug_report_bug_report__fs-oracleCommonVoice17_validated_faSWE-bench_validated_12_18_unsolved_style-3__fs-oraclelandsea-validated-leadsSWE-bench_validated_12_22_style-3__fs-oraclecube_validated_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "",
"total_episodes": 91,
"total_frames": 115522,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:91"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/gauravpradeep/cube_validated_lerobot.cube_validatedimabari_wiki_qa_v4_validated
Imabari QA v4 — Validated
Dataset Summary
Imabari QA v4 — Validated is a Japanese question-answering dataset for supervised fine-tuning (SFT).
This dataset combines two independently curated variants of the Imabari QA v4 synthetic reasoning dataset:
ikedachin/imabari_qa_v4_program_validated
ikedachin/imabari_qa_v4_human_validated
Both datasets are derived from:
Source corpus: ikedachin/imabari_wiki_cpt_v3
QA generation model: Qwen3.8-27B-NVFP4
The combined… See the full description on the dataset page: https://huggingface.co/datasets/ikedachin/imabari_wiki_qa_v4_validated.ML-1M-Syntax-Validated-Python-Code
ML-1M Syntax-Validated Python Code
Dataset Summary
ML-1M Syntax-Validated Python Code is a large-scale corpus containing over 1 million machine-learning–oriented Python programs derived from The Stack, a permissively licensed collection of open-source source code.
The dataset is constructed through heuristic ML-domain filtering, syntactic validation, and basic safety checks. It is intended to support empirical analysis of real-world ML code, executability and dependency… See the full description on the dataset page: https://huggingface.co/datasets/Noushad999/ML-1M-Syntax-Validated-Python-Code.hpca2027-database-validated-traces-20260711
HPCA 2027 validated database traces
Validated DynamoRIO instruction traces for database application-plus-input
configurations used in HPCA 2027 backend-bound characterization.
Included workloads
Directory
Application and input
Captured fetched instructions
Scarab backend bound
mongodb
MongoDB 7.0.37, 10M-record YCSB database, sort/aggregate, 1 GiB WiredTiger cache, 3 GiB container
100,098,204 global across server threads
41.76% on the 89,501… See the full description on the dataset page: https://huggingface.co/datasets/harry1332/hpca2027-database-validated-traces-20260711.synthvision-validated-qwen-by-kimi
synthvision-validated-qwen-by-kimi
Qwen 3.5 annotations validated by Kimi K2.5 (93.1% pass rate)
Records: 55,359
About
Cross-validated subset from the SynthVision pipeline. Kimi K2.5 reviewed all 59,476 Qwen 3.5 annotations and confirmed 55,359 as consistent with the source images (93.1% pass rate).
Validation criteria: consistent == true AND confidence >= 0.7. Records that failed validation were removed — primarily cases where the annotator hallucinated findings not… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/synthvision-validated-qwen-by-kimi.5_synt_flux_landmark_selected_single_validated_1011robotsim-soma-source-shards-validated-v01auto_validated2hpca2027-agentic-validated-traces-20260711
HPCA 2027 CPU-Local Agentic Validated Traces
This dataset contains validated CPU-local traces used for HPCA 2027 agentic
workload characterization. AppWorld, CORE-Bench, and Terminal-Bench are
paper-backed benchmark workloads from official repositories; the hybrid-RAG
and data-analysis agents are deterministic characterization workloads retained
from the initial pass.
Workloads And Full 100M Results
Workload
Source
Frontend
Backend
Bad speculation
Retiring… See the full description on the dataset page: https://huggingface.co/datasets/harry1332/hpca2027-agentic-validated-traces-20260711.syn_validatedauto_validatedrunyankole-crowd-validated-paths
Dataset Card for "runyankole-crowd-validated-paths"
More Information needed
validated_12_18bitflags_validatedregex_validatedvalidated_12_18_unsolved
