datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhysicalAI-SimReady-Warehouse-01
NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset
Dataset Version: 1.1.0
Date: May 18, 2025
Author: NVIDIA, Corporation
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Contents
This dataset includes the following:
This README file
A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is organized in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-SimReady-Warehouse-01.quantum-representations
Epsilon-Transformers Belief Analysis Dataset
This dataset contains trained neural network models and their corresponding belief state regression analysis from the Epsilon-Transformers project. The models were trained on four different stochastic processes and analyzed for their ability to learn and represent belief states.
See https://github.com/adamimos/epsilon-transformers/tree/quantum-public for codebase which generated this data.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/SimplexAI/quantum-representations.python-fim
Python Stack | Fill-in-the-Middle
This is a conversion or adaptation of The Stack to a python FIM task. The example column is B64 encoded because people like to put special characters in their code that csv files dont like so I encoded the strings before saving them to disk.
SimpleQA
SimpleQA
A factuality benchmark called SimpleQA that measures the ability for language models to answer short, fact-seeking questions.
Sources
openai/simple-evals
Introducing SimpleQA
Measuring short-form factuality in large language models
simpleqa-verified
SimpleQA Verified
A 1,000-prompt factuality benchmark from Google DeepMind and Google Research, designed to reliably evaluate LLM parametric knowledge.
▶ SimpleQA Verified Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code
Benchmark
SimpleQA Verified is a 1,000-prompt benchmark for reliably evaluating Large Language Models (LLMs) on short-form factuality
and parametric knowledge. The authors from Google DeepMind and Google Research… See the full description on the dataset page: https://huggingface.co/datasets/google/simpleqa-verified.SimpleSafetyTestssim-datasets
SIM-Datasets: A Unified Symbolic Regression Benchmark
A standardized benchmark collection designed for the Scientific Intelligent Modelling (SIM) toolkit, providing comprehensive datasets for symbolic regression research and applications.
Overview
SIM-Datasets serves as a unified benchmark for symbolic regression tasks, offering standardized datasets with consistent formatting and evaluation protocols. This collection is specifically curated to support the Scientific… See the full description on the dataset page: https://huggingface.co/datasets/scientific-intelligent-modelling/sim-datasets.simplified_grooveThis is a copy of the Magenta Groove dataset
The script ´simplify_midi_pretty.py` reads the midi data and simplifies it, by removing any midi values that aren't kicks or snares, and quantizing the notes.
SimpleQA-VerifiedSimpleQA Verified is a 1,000-prompt benchmark for reliably evaluating Large Language Models (LLMs) on short-form factuality and parametric knowledge. The authors from Google DeepMind and Google Research address various limitations of SimpleQA, originally designed by Wei et al. (2024) at OpenAI, including noisy and incorrect labels, topical biases, and question redundancy. SimpleQA Verified was created to provide the research community with a more precise instrument to track genuine progress in… See the full description on the dataset page: https://huggingface.co/datasets/codelion/SimpleQA-Verified.europarl_for_language_detection_10kSimpleQA-VerifiedSimpleQA Verified is a 1,000-prompt benchmark for reliably evaluating Large Language Models (LLMs) on short-form factuality and parametric knowledge. The authors from Google DeepMind and Google Research address various limitations of SimpleQA, originally designed by Wei et al. (2024) at OpenAI, including noisy and incorrect labels, topical biases, and question redundancy.
SimpleQA Verified was created to provide the research community with a more precise instrument to track genuine progress in… See the full description on the dataset page: https://huggingface.co/datasets/stalkermustang/SimpleQA-Verified.3d-models-for-isaac-sim-dataset
Dataset of 3D models for Isaac Sim (USDZ)
🇬🇧 English Description
This dataset contains a collection of 3D models converted to the .usdz format, featuring proper Semantic Labeling. These assets are optimized for generating synthetic training data using NVIDIA Isaac Sim and NVIDIA Replicator.
Primary Use Case: Training object detection and segmentation models (e.g., YOLO, RT-DETR, Mask R-CNN).
Class List
The dataset includes the following 30 semantic… See the full description on the dataset page: https://huggingface.co/datasets/barszot/3d-models-for-isaac-sim-dataset.LM-SimBench
LM-SimBench
Dataset Description
LM-SimBench is a large-scale training-performance profiling dataset for large language models. The dataset is collected from training runs based on the MindSpeed-LLM framework and the Ascend NPU development stack, covering multiple model families, context lengths, and distributed parallel configurations.
Each model is sampled under feasible combinations of data parallelism (DP), tensor parallelism (TP), pipeline parallelism (PP), context… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-neurips26-ljasd/LM-SimBench.PhysicalAI-SimReady-Warehouse-01
NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset
Dataset Version: 1.1.0
Date: May 18, 2025
Author: NVIDIA, Corporation
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Contents
This dataset includes the following:
This README file
A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is… See the full description on the dataset page: https://huggingface.co/datasets/yzhllm/PhysicalAI-SimReady-Warehouse-01.ising-sim2real-results
Ising sim2real — Decoder Benchmark Results
Evaluation results for a panel of open surface-code decoders run on real Google
Willow hardware data and on synthetic circuit-level noise of rising fidelity.
The question these results answer: does the cheap synthetic benchmark predict the
real-hardware result?
This repo holds the outputs (LERs, per-shot outcomes, fitted noise models,
figures). The inputs — ingested Willow detection events, circuits, and shipped
DEMs — live in… See the full description on the dataset page: https://huggingface.co/datasets/ShayManor/ising-sim2real-results.gnss-nmea-jammertest25
GNSS NMEA Jammertest 2025
Raw receiver telemetry collected during the Jammertest 2025 campaign, used in the paper "GNSS Jamming and Spoofing Detection Using NMEA Data" (ICL-GNSS 2026).
Companion code (preprocessing, feature engineering, model training/evaluation) is available at: https://github.com/simula/icl-gnss26
Data Collection
The data was collected over five days in September 2025 on Andøya Island, Norway, across three test sites. The experimental plan is… See the full description on the dataset page: https://huggingface.co/datasets/SimulaMet/gnss-nmea-jammertest25.simlexJailbreakPrompts
Independent Jailbreak Datasets for LLM Guardrail Evaluation
Constructed for the thesis:“Contamination Effects: How Training Data Leakage Affects Red Team Evaluation of LLM Jailbreak Detection”
The effectiveness of LLM guardrails is commonly evaluated using open-source red teaming tools. However, this study reveals that significant data contamination exists between the training sets of binary jailbreak classifiers (ProtectAI, Katanemo, TestSavantAI, etc.) and the test prompts used in… See the full description on the dataset page: https://huggingface.co/datasets/Simsonsun/JailbreakPrompts.textclassificationMNLIamazon_zh_simpleLM-SimBench_example
LM-SimBench (Example Snapshot)
Dataset Description
This repository distributes a compact example snapshot of LM-SimBench, the structured CSV release of large-scale LLM training-performance profiling data. The snapshot is provided so reviewers and readers can inspect file layout, schemas, and representative records without downloading the multi–tens-of-gigabyte full release.
The profiling methodology, software stack, and field definitions are the same as in the complete… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-neurips26-ljasd/LM-SimBench_example.AIDE-Chip-15K-gem5-Sims
AIDE-Chip 15K gem5 Simulation Dataset
AIDE-Chip-15K-gem5-Sims is a structured dataset of approximately 15,000 validated RISC-V gem5 simulations covering cache hierarchy design-space exploration (DSE) for single-core processors.
The dataset was generated using gem5's Syscall Emulation (SE) mode and six representative workloads, spanning compute-bound, memory-bound, and irregular access patterns. Each sample maps cache configuration parameters to IPC and L2 miss rate, enabling… See the full description on the dataset page: https://huggingface.co/datasets/uralstech/AIDE-Chip-15K-gem5-Sims.wikipedia-22-12-simple-embeddings
wikipedia-22-12-simple-embeddings
A modified version of Cohere/wikipedia-22-12-simple-embeddings
meant for use with PostgreSQL with pgvector and Timescale Vector.
Dataset Details
This dataset was created for exploring time-based filtering and semantic search in PostgreSQL with pgvector and Timescale Vector.
This is a modified version of the Cohere wikipedia-22-12-simple-embeddings dataset hosted on Huggingface.
It contains embeddings of Simple English Wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/timescale/wikipedia-22-12-simple-embeddings.SimpleToolCallingchar-sim-data
Dataset Card for Character Similarity Dataset
Dataset Details
The Character Similarity Dataset is a collection of textual trait descriptions along with the corresponding ontology based similarity measures between trait description pairs. The distance is estimated using the Phenoscape Knowledgebase as the ontology. The Knowledgebase is built upon a number of OBO ontologies, most importantly the Uberon anatomy ontology.
The Character Similarity Dataset is a collection of… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/char-sim-data.SimpleMath
🧮 SimpleMath 100K
SimpleMath 100K is a high-quality synthetic dataset of 100,000 basic arithmetic problems — no noise, no tricks, just clean and accurate math.
✅ Purpose
This was made for small AI models — not to struggle with complex math, but to get simple math right every time.
📦 Contents
75,000 numeric problems, evenly split:
18,750 addition (456 + 789 =)
18,750 subtraction (900 - 345 =)
18,750 multiplication (12 x 15 =)
18,750 division (144 / 12 =)… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/SimpleMath.simplifyweibo_4_moodsdkt-dataset-simFarm_Sim_Dataspan-similarity-dataset
Span Similarity Dataset (SSD)
Dataset Summary
The Span Similarity Dataset (SSD) focuses on Explainable Textual Similarity. It consists
of pairs of sentences with annotations pointing to both semantically equivalent and
dissimilar spans.
Languages
The SSD includes exclusively texts in English.
Dataset Structure
The dataset is split into -train (800 samples), -eval (100 samples), and -test (100
samples), all of them provided as a .tsv… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/span-similarity-dataset.
