CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /PhysicalAI-SimReady-Warehouse-01 NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset Dataset Version: 1.1.0 Date: May 18, 2025 Author: NVIDIA, Corporation License: CC-BY-4.0 (Creative Commons Attribution 4.0 International) Contents This dataset includes the following: This README file A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is organized in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-SimReady-Warehouse-01.imageimage-segmentationn<1K54 likes27k downloads10mo agoHugging Face02simmo /python-fim Python Stack | Fill-in-the-Middle This is a conversion or adaptation of The Stack to a python FIM task. The example column is B64 encoded because people like to put special characters in their code that csv files dont like so I encoded the strings before saving them to disk. textfill-mask10M<n<100M0 likes4.6k downloads2y agoHugging Face03basicv8vc /SimpleQA SimpleQA A factuality benchmark called SimpleQA that measures the ability for language models to answer short, fact-seeking questions. Sources openai/simple-evals Introducing SimpleQA Measuring short-form factuality in large language models textquestion-answering1K<n<10K32 likes3.6k downloads2y agoHugging Face04google /simpleqa-verified SimpleQA Verified A 1,000-prompt factuality benchmark from Google DeepMind and Google Research, designed to reliably evaluate LLM parametric knowledge. ▶ SimpleQA Verified Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code Benchmark SimpleQA Verified is a 1,000-prompt benchmark for reliably evaluating Large Language Models (LLMs) on short-form factuality and parametric knowledge. The authors from Google DeepMind and Google Research… See the full description on the dataset page: https://huggingface.co/datasets/google/simpleqa-verified.textquestion-answering1K<n<10K53 likes3.5k downloads7mo agoHugging Face05Bertievidgen /SimpleSafetyTeststexttext-generationn<1K12 likes3.1k downloads2y agoHugging Face06MidiAndTheGang /simplified_grooveThis is a copy of the Magenta Groove dataset The script ´simplify_midi_pretty.py` reads the midi data and simplifies it, by removing any midi values that aren't kicks or snares, and quantizing the notes. tabular1K<n<10K0 likes2.2k downloads2y agoHugging Face07codelion /SimpleQA-VerifiedSimpleQA Verified is a 1,000-prompt benchmark for reliably evaluating Large Language Models (LLMs) on short-form factuality and parametric knowledge. The authors from Google DeepMind and Google Research address various limitations of SimpleQA, originally designed by Wei et al. (2024) at OpenAI, including noisy and incorrect labels, topical biases, and question redundancy. SimpleQA Verified was created to provide the research community with a more precise instrument to track genuine progress in… See the full description on the dataset page: https://huggingface.co/datasets/codelion/SimpleQA-Verified.text1K<n<10K4 likes1.9k downloads1y agoHugging Face08simoneteglia /europarl_for_language_detection_10ktext100K<n<1M0 likes739 downloads3y agoHugging Face09stalkermustang /SimpleQA-VerifiedSimpleQA Verified is a 1,000-prompt benchmark for reliably evaluating Large Language Models (LLMs) on short-form factuality and parametric knowledge. The authors from Google DeepMind and Google Research address various limitations of SimpleQA, originally designed by Wei et al. (2024) at OpenAI, including noisy and incorrect labels, topical biases, and question redundancy. SimpleQA Verified was created to provide the research community with a more precise instrument to track genuine progress in… See the full description on the dataset page: https://huggingface.co/datasets/stalkermustang/SimpleQA-Verified.textquestion-answering1K<n<10K0 likes560 downloads1y agoHugging Face10barszot /3d-models-for-isaac-sim-dataset Dataset of 3D models for Isaac Sim (USDZ) 🇬🇧 English Description This dataset contains a collection of 3D models converted to the .usdz format, featuring proper Semantic Labeling. These assets are optimized for generating synthetic training data using NVIDIA Isaac Sim and NVIDIA Replicator. Primary Use Case: Training object detection and segmentation models (e.g., YOLO, RT-DETR, Mask R-CNN). Class List The dataset includes the following 30 semantic… See the full description on the dataset page: https://huggingface.co/datasets/barszot/3d-models-for-isaac-sim-dataset.3dn<1K0 likes505 downloads6mo agoHugging Face11anonymous-neurips26-ljasd /LM-SimBench LM-SimBench Dataset Description LM-SimBench is a large-scale training-performance profiling dataset for large language models. The dataset is collected from training runs based on the MindSpeed-LLM framework and the Ascend NPU development stack, covering multiple model families, context lengths, and distributed parallel configurations. Each model is sampled under feasible combinations of data parallelism (DP), tensor parallelism (TP), pipeline parallelism (PP), context… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-neurips26-ljasd/LM-SimBench.tabulartabular-regressionn<1K0 likes466 downloads5mo agoHugging Face12ShayManor /ising-sim2real-results Ising sim2real — Decoder Benchmark Results Evaluation results for a panel of open surface-code decoders run on real Google Willow hardware data and on synthetic circuit-level noise of rising fidelity. The question these results answer: does the cheap synthetic benchmark predict the real-hardware result? This repo holds the outputs (LERs, per-shot outcomes, fitted noise models, figures). The inputs — ingested Willow detection events, circuits, and shipped DEMs — live in… See the full description on the dataset page: https://huggingface.co/datasets/ShayManor/ising-sim2real-results.tabularother10K<n<100K1 likes414 downloads12d agoHugging Face13yzhllm /PhysicalAI-SimReady-Warehouse-01 NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset Dataset Version: 1.1.0 Date: May 18, 2025 Author: NVIDIA, Corporation License: CC-BY-4.0 (Creative Commons Attribution 4.0 International) Contents This dataset includes the following: This README file A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is… See the full description on the dataset page: https://huggingface.co/datasets/yzhllm/PhysicalAI-SimReady-Warehouse-01.imageimage-segmentationn<1K0 likes376 downloads4mo agoHugging Face14SimulaMet /gnss-nmea-jammertest25 GNSS NMEA Jammertest 2025 Raw receiver telemetry collected during the Jammertest 2025 campaign, used in the paper "GNSS Jamming and Spoofing Detection Using NMEA Data" (ICL-GNSS 2026). Companion code (preprocessing, feature engineering, model training/evaluation) is available at: https://github.com/simula/icl-gnss26 Data Collection The data was collected over five days in September 2025 on Andøya Island, Norway, across three test sites. The experimental plan is… See the full description on the dataset page: https://huggingface.co/datasets/SimulaMet/gnss-nmea-jammertest25.tabulartabular-classification1M<n<10M1 likes312 downloads1mo agoHugging Face15tasksource /simlextabularn<1K0 likes281 downloads3y agoHugging Face16Simsonsun /JailbreakPrompts Independent Jailbreak Datasets for LLM Guardrail Evaluation Constructed for the thesis:“Contamination Effects: How Training Data Leakage Affects Red Team Evaluation of LLM Jailbreak Detection” The effectiveness of LLM guardrails is commonly evaluated using open-source red teaming tools. However, this study reveals that significant data contamination exists between the training sets of binary jailbreak classifiers (ProtectAI, Katanemo, TestSavantAI, etc.) and the test prompts used in… See the full description on the dataset page: https://huggingface.co/datasets/Simsonsun/JailbreakPrompts.text1K<n<10K13 likes264 downloads1y agoHugging Face17simana /textclassificationMNLItext100K<n<1M0 likes235 downloads4y agoHugging Face18ASCCCCCCCC /amazon_zh_simpletext10K<n<100K1 likes180 downloads5y agoHugging Face19anonymous-neurips26-ljasd /LM-SimBench_example LM-SimBench (Example Snapshot) Dataset Description This repository distributes a compact example snapshot of LM-SimBench, the structured CSV release of large-scale LLM training-performance profiling data. The snapshot is provided so reviewers and readers can inspect file layout, schemas, and representative records without downloading the multi–tens-of-gigabyte full release. The profiling methodology, software stack, and field definitions are the same as in the complete… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-neurips26-ljasd/LM-SimBench_example.texttabular-regressionn<1K0 likes178 downloads5mo agoHugging Face20uralstech /AIDE-Chip-15K-gem5-Sims AIDE-Chip 15K gem5 Simulation Dataset AIDE-Chip-15K-gem5-Sims is a structured dataset of approximately 15,000 validated RISC-V gem5 simulations covering cache hierarchy design-space exploration (DSE) for single-core processors. The dataset was generated using gem5's Syscall Emulation (SE) mode and six representative workloads, spanning compute-bound, memory-bound, and irregular access patterns. Each sample maps cache configuration parameters to IPC and L2 miss rate, enabling… See the full description on the dataset page: https://huggingface.co/datasets/uralstech/AIDE-Chip-15K-gem5-Sims.tabulartabular-regression10K<n<100K0 likes161 downloads8mo agoHugging Face21timescale /wikipedia-22-12-simple-embeddings wikipedia-22-12-simple-embeddings A modified version of Cohere/wikipedia-22-12-simple-embeddings meant for use with PostgreSQL with pgvector and Timescale Vector. Dataset Details This dataset was created for exploring time-based filtering and semantic search in PostgreSQL with pgvector and Timescale Vector. This is a modified version of the Cohere wikipedia-22-12-simple-embeddings dataset hosted on Huggingface. It contains embeddings of Simple English Wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/timescale/wikipedia-22-12-simple-embeddings.texttext-retrieval100K<n<1M0 likes154 downloads3y agoHugging Face22bebechien /SimpleToolCallingtextn<1K4 likes141 downloads10mo agoHugging Face23imageomics /char-sim-data Dataset Card for Character Similarity Dataset Dataset Details The Character Similarity Dataset is a collection of textual trait descriptions along with the corresponding ontology based similarity measures between trait description pairs. The distance is estimated using the Phenoscape Knowledgebase as the ontology. The Knowledgebase is built upon a number of OBO ontologies, most importantly the Uberon anatomy ontology. The Character Similarity Dataset is a collection of… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/char-sim-data.tabularfeature-extraction100M<n<1B0 likes127 downloads11mo agoHugging Face24dirtycomputer /simplifyweibo_4_moodstext100K<n<1M4 likes115 downloads4y agoHugging Face25ProCreations /SimpleMath 🧮 SimpleMath 100K SimpleMath 100K is a high-quality synthetic dataset of 100,000 basic arithmetic problems — no noise, no tricks, just clean and accurate math. ✅ Purpose This was made for small AI models — not to struggle with complex math, but to get simple math right every time. 📦 Contents 75,000 numeric problems, evenly split: 18,750 addition (456 + 789 =) 18,750 subtraction (900 - 345 =) 18,750 multiplication (12 x 15 =) 18,750 division (144 / 12 =)… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/SimpleMath.texttext-generation100K<n<1M8 likes110 downloads1y agoHugging Face26UI-Simulator /UI-Simulator_web_datatext10K<n<100K1 likes105 downloads1y agoHugging Face27UI-Simulator /UI-Simulator_android_datatext10K<n<100K3 likes97 downloads1y agoHugging Face28tum-nlp /span-similarity-dataset Span Similarity Dataset (SSD) Dataset Summary The Span Similarity Dataset (SSD) focuses on Explainable Textual Similarity. It consists of pairs of sentences with annotations pointing to both semantically equivalent and dissimilar spans. Languages The SSD includes exclusively texts in English. Dataset Structure The dataset is split into -train (800 samples), -eval (100 samples), and -test (100 samples), all of them provided as a .tsv… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/span-similarity-dataset.textsentence-similarity1K<n<10K2 likes93 downloads2mo agoHugging Face29r1char9 /simplification-datasetДанный dataset был собран из корпуса "RuSimpleSentEval" (https://github.com/dialogue-evaluation/RuSimpleSentEval), а также "RuAdapt" (https://github.com/Digital-Pushkin-Lab/RuAdapt) для задачи упрощения текста (text simplification). from datasets import load_dataset data_files = {'train':"train.csv",'test':"test.csv"} dataset = load_dataset("r1char9/simplification", data_files=data_files) train_df = dataset['train'].to_pandas() test_df = dataset['test'].to_pandas() text1K<n<10K0 likes91 downloads24d agoHugging Face30dongko /dkt-dataset-simtabular100K<n<1M0 likes90 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.