CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /PhysicalAI-SimReady-Warehouse-01 NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset Dataset Version: 1.1.0 Date: May 18, 2025 Author: NVIDIA, Corporation License: CC-BY-4.0 (Creative Commons Attribution 4.0 International) Contents This dataset includes the following: This README file A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is organized in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-SimReady-Warehouse-01.imageimage-segmentationn<1K53 likes27k downloads10mo agoHugging Face02SimplexAI /quantum-representations Epsilon-Transformers Belief Analysis Dataset This dataset contains trained neural network models and their corresponding belief state regression analysis from the Epsilon-Transformers project. The models were trained on four different stochastic processes and analyzed for their ability to learn and represent belief states. See https://github.com/adamimos/epsilon-transformers/tree/quantum-public for codebase which generated this data. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/SimplexAI/quantum-representations.tabularother100K<n<1M0 likes6.1k downloads1y agoHugging Face03simmo /python-fim Python Stack | Fill-in-the-Middle This is a conversion or adaptation of The Stack to a python FIM task. The example column is B64 encoded because people like to put special characters in their code that csv files dont like so I encoded the strings before saving them to disk. textfill-mask10M<n<100M0 likes4.5k downloads2y agoHugging Face04basicv8vc /SimpleQA SimpleQA A factuality benchmark called SimpleQA that measures the ability for language models to answer short, fact-seeking questions. Sources openai/simple-evals Introducing SimpleQA Measuring short-form factuality in large language models textquestion-answering1K<n<10K32 likes3.6k downloads2y agoHugging Face05google /simpleqa-verified SimpleQA Verified A 1,000-prompt factuality benchmark from Google DeepMind and Google Research, designed to reliably evaluate LLM parametric knowledge. ▶ SimpleQA Verified Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code Benchmark SimpleQA Verified is a 1,000-prompt benchmark for reliably evaluating Large Language Models (LLMs) on short-form factuality and parametric knowledge. The authors from Google DeepMind and Google Research… See the full description on the dataset page: https://huggingface.co/datasets/google/simpleqa-verified.textquestion-answering1K<n<10K52 likes3.4k downloads7mo agoHugging Face06Bertievidgen /SimpleSafetyTeststexttext-generationn<1K12 likes3.1k downloads2y agoHugging Face07scientific-intelligent-modelling /sim-datasets SIM-Datasets: A Unified Symbolic Regression Benchmark A standardized benchmark collection designed for the Scientific Intelligent Modelling (SIM) toolkit, providing comprehensive datasets for symbolic regression research and applications. Overview SIM-Datasets serves as a unified benchmark for symbolic regression tasks, offering standardized datasets with consistent formatting and evaluation protocols. This collection is specifically curated to support the Scientific… See the full description on the dataset page: https://huggingface.co/datasets/scientific-intelligent-modelling/sim-datasets.tabular10M<n<100M0 likes2.7k downloads1y agoHugging Face08MidiAndTheGang /simplified_grooveThis is a copy of the Magenta Groove dataset The script ´simplify_midi_pretty.py` reads the midi data and simplifies it, by removing any midi values that aren't kicks or snares, and quantizing the notes. tabular1K<n<10K0 likes2.1k downloads2y agoHugging Face09codelion /SimpleQA-VerifiedSimpleQA Verified is a 1,000-prompt benchmark for reliably evaluating Large Language Models (LLMs) on short-form factuality and parametric knowledge. The authors from Google DeepMind and Google Research address various limitations of SimpleQA, originally designed by Wei et al. (2024) at OpenAI, including noisy and incorrect labels, topical biases, and question redundancy. SimpleQA Verified was created to provide the research community with a more precise instrument to track genuine progress in… See the full description on the dataset page: https://huggingface.co/datasets/codelion/SimpleQA-Verified.text1K<n<10K4 likes1.8k downloads1y agoHugging Face10simoneteglia /europarl_for_language_detection_10ktext100K<n<1M0 likes649 downloads3y agoHugging Face11stalkermustang /SimpleQA-VerifiedSimpleQA Verified is a 1,000-prompt benchmark for reliably evaluating Large Language Models (LLMs) on short-form factuality and parametric knowledge. The authors from Google DeepMind and Google Research address various limitations of SimpleQA, originally designed by Wei et al. (2024) at OpenAI, including noisy and incorrect labels, topical biases, and question redundancy. SimpleQA Verified was created to provide the research community with a more precise instrument to track genuine progress in… See the full description on the dataset page: https://huggingface.co/datasets/stalkermustang/SimpleQA-Verified.textquestion-answering1K<n<10K0 likes579 downloads1y agoHugging Face12barszot /3d-models-for-isaac-sim-dataset Dataset of 3D models for Isaac Sim (USDZ) 🇬🇧 English Description This dataset contains a collection of 3D models converted to the .usdz format, featuring proper Semantic Labeling. These assets are optimized for generating synthetic training data using NVIDIA Isaac Sim and NVIDIA Replicator. Primary Use Case: Training object detection and segmentation models (e.g., YOLO, RT-DETR, Mask R-CNN). Class List The dataset includes the following 30 semantic… See the full description on the dataset page: https://huggingface.co/datasets/barszot/3d-models-for-isaac-sim-dataset.3dn<1K0 likes508 downloads6mo agoHugging Face13anonymous-neurips26-ljasd /LM-SimBench LM-SimBench Dataset Description LM-SimBench is a large-scale training-performance profiling dataset for large language models. The dataset is collected from training runs based on the MindSpeed-LLM framework and the Ascend NPU development stack, covering multiple model families, context lengths, and distributed parallel configurations. Each model is sampled under feasible combinations of data parallelism (DP), tensor parallelism (TP), pipeline parallelism (PP), context… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-neurips26-ljasd/LM-SimBench.tabulartabular-regressionn<1K0 likes467 downloads5mo agoHugging Face14yzhllm /PhysicalAI-SimReady-Warehouse-01 NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset Dataset Version: 1.1.0 Date: May 18, 2025 Author: NVIDIA, Corporation License: CC-BY-4.0 (Creative Commons Attribution 4.0 International) Contents This dataset includes the following: This README file A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is… See the full description on the dataset page: https://huggingface.co/datasets/yzhllm/PhysicalAI-SimReady-Warehouse-01.imageimage-segmentationn<1K0 likes368 downloads4mo agoHugging Face15ShayManor /ising-sim2real-results Ising sim2real — Decoder Benchmark Results Evaluation results for a panel of open surface-code decoders run on real Google Willow hardware data and on synthetic circuit-level noise of rising fidelity. The question these results answer: does the cheap synthetic benchmark predict the real-hardware result? This repo holds the outputs (LERs, per-shot outcomes, fitted noise models, figures). The inputs — ingested Willow detection events, circuits, and shipped DEMs — live in… See the full description on the dataset page: https://huggingface.co/datasets/ShayManor/ising-sim2real-results.tabularother10K<n<100K1 likes344 downloads11d agoHugging Face16SimulaMet /gnss-nmea-jammertest25 GNSS NMEA Jammertest 2025 Raw receiver telemetry collected during the Jammertest 2025 campaign, used in the paper "GNSS Jamming and Spoofing Detection Using NMEA Data" (ICL-GNSS 2026). Companion code (preprocessing, feature engineering, model training/evaluation) is available at: https://github.com/simula/icl-gnss26 Data Collection The data was collected over five days in September 2025 on Andøya Island, Norway, across three test sites. The experimental plan is… See the full description on the dataset page: https://huggingface.co/datasets/SimulaMet/gnss-nmea-jammertest25.tabulartabular-classification1M<n<10M1 likes314 downloads1mo agoHugging Face17tasksource /simlextabularn<1K0 likes287 downloads3y agoHugging Face18Simsonsun /JailbreakPrompts Independent Jailbreak Datasets for LLM Guardrail Evaluation Constructed for the thesis:“Contamination Effects: How Training Data Leakage Affects Red Team Evaluation of LLM Jailbreak Detection” The effectiveness of LLM guardrails is commonly evaluated using open-source red teaming tools. However, this study reveals that significant data contamination exists between the training sets of binary jailbreak classifiers (ProtectAI, Katanemo, TestSavantAI, etc.) and the test prompts used in… See the full description on the dataset page: https://huggingface.co/datasets/Simsonsun/JailbreakPrompts.text1K<n<10K12 likes260 downloads1y agoHugging Face19simana /textclassificationMNLItext100K<n<1M0 likes221 downloads4y agoHugging Face20ASCCCCCCCC /amazon_zh_simpletext10K<n<100K1 likes178 downloads5y agoHugging Face21anonymous-neurips26-ljasd /LM-SimBench_example LM-SimBench (Example Snapshot) Dataset Description This repository distributes a compact example snapshot of LM-SimBench, the structured CSV release of large-scale LLM training-performance profiling data. The snapshot is provided so reviewers and readers can inspect file layout, schemas, and representative records without downloading the multi–tens-of-gigabyte full release. The profiling methodology, software stack, and field definitions are the same as in the complete… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-neurips26-ljasd/LM-SimBench_example.texttabular-regressionn<1K0 likes177 downloads5mo agoHugging Face22uralstech /AIDE-Chip-15K-gem5-Sims AIDE-Chip 15K gem5 Simulation Dataset AIDE-Chip-15K-gem5-Sims is a structured dataset of approximately 15,000 validated RISC-V gem5 simulations covering cache hierarchy design-space exploration (DSE) for single-core processors. The dataset was generated using gem5's Syscall Emulation (SE) mode and six representative workloads, spanning compute-bound, memory-bound, and irregular access patterns. Each sample maps cache configuration parameters to IPC and L2 miss rate, enabling… See the full description on the dataset page: https://huggingface.co/datasets/uralstech/AIDE-Chip-15K-gem5-Sims.tabulartabular-regression10K<n<100K0 likes162 downloads8mo agoHugging Face23timescale /wikipedia-22-12-simple-embeddings wikipedia-22-12-simple-embeddings A modified version of Cohere/wikipedia-22-12-simple-embeddings meant for use with PostgreSQL with pgvector and Timescale Vector. Dataset Details This dataset was created for exploring time-based filtering and semantic search in PostgreSQL with pgvector and Timescale Vector. This is a modified version of the Cohere wikipedia-22-12-simple-embeddings dataset hosted on Huggingface. It contains embeddings of Simple English Wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/timescale/wikipedia-22-12-simple-embeddings.texttext-retrieval100K<n<1M0 likes154 downloads3y agoHugging Face24bebechien /SimpleToolCallingtextn<1K4 likes140 downloads10mo agoHugging Face25imageomics /char-sim-data Dataset Card for Character Similarity Dataset Dataset Details The Character Similarity Dataset is a collection of textual trait descriptions along with the corresponding ontology based similarity measures between trait description pairs. The distance is estimated using the Phenoscape Knowledgebase as the ontology. The Knowledgebase is built upon a number of OBO ontologies, most importantly the Uberon anatomy ontology. The Character Similarity Dataset is a collection of… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/char-sim-data.tabularfeature-extraction100M<n<1B0 likes126 downloads11mo agoHugging Face26ProCreations /SimpleMath 🧮 SimpleMath 100K SimpleMath 100K is a high-quality synthetic dataset of 100,000 basic arithmetic problems — no noise, no tricks, just clean and accurate math. ✅ Purpose This was made for small AI models — not to struggle with complex math, but to get simple math right every time. 📦 Contents 75,000 numeric problems, evenly split: 18,750 addition (456 + 789 =) 18,750 subtraction (900 - 345 =) 18,750 multiplication (12 x 15 =) 18,750 division (144 / 12 =)… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/SimpleMath.texttext-generation100K<n<1M8 likes117 downloads1y agoHugging Face27dirtycomputer /simplifyweibo_4_moodstext100K<n<1M4 likes113 downloads4y agoHugging Face28dongko /dkt-dataset-simtabular100K<n<1M0 likes107 downloads2y agoHugging Face29jeffrey1963 /Farm_Sim_Datatabularn<1K0 likes99 downloads1y agoHugging Face30tum-nlp /span-similarity-dataset Span Similarity Dataset (SSD) Dataset Summary The Span Similarity Dataset (SSD) focuses on Explainable Textual Similarity. It consists of pairs of sentences with annotations pointing to both semantically equivalent and dissimilar spans. Languages The SSD includes exclusively texts in English. Dataset Structure The dataset is split into -train (800 samples), -eval (100 samples), and -test (100 samples), all of them provided as a .tsv… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/span-similarity-dataset.textsentence-similarity1K<n<10K2 likes93 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.