CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01datasets-maintainers /dataset-with-standalone-yamlThis is a test dataset used in the datasets library CI textn<1K0 likes18k downloads3y agoHugging Face02mjuicem /StreamingBench StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding 🏠 Project Page | 📄 arXiv Paper | 📦 Dataset | 🏅Leaderboard StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟 [NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output. [NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/mjuicem/StreamingBench.imagequestion-answering1K<n<10K13 likes12k downloads1y agoHugging Face03stanford-crfm /air-bench-2024 AIRBench 2024 AIRBench 2024 is a AI safety benchmark that aligns with emerging government regulations and company policies. It consists of diverse, malicious prompts spanning categories of the regulation-based safety categories in the AIR 2024 safety taxonomy. Dataset Details Dataset Description AIRBench 2024 is a AI safety benchmark that aligns with emerging government regulations and company policies. It consists of diverse, malicious prompts spanning… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/air-bench-2024.texttext-generation10K<n<100K26 likes7.6k downloads2y agoHugging Face04prquan /STARK_10k STARK: Spatial-Temporal reAsoning benchmaRK STARK is a comprehensive benchmark designed to systematically evaluate large language models (LLMs) and large reasoning models (LRMs) on spatial-temporal reasoning tasks, particularly for applications in cyber-physical systems (CPS) such as robotics, autonomous vehicles, and smart city infrastructure. Dataset Summary Hierarchical Benchmark: Tasks are structured across three levels of reasoning complexity: State Estimation:… See the full description on the dataset page: https://huggingface.co/datasets/prquan/STARK_10k.textquestion-answering10K<n<100K1 likes7k downloads11mo agoHugging Face05stdKonjac /Sparkle Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance Ziyun Zeng, Yiqi Lin, Guoqiang Liang, and Mike Zheng Shou 📦 Dataset Sparkle is a large-scale video background replacement dataset comprising ~140K high-quality source–edited video pairs. It is fully open-sourced at 🤗stdKonjac/Sparkle. For full methodology and dataset details, please refer to our paper. The dataset is organized into five themes along different… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/Sparkle.imagetext-to-video100K<n<1M1 likes4.4k downloads5mo agoHugging Face06CTPLab-DBE-UniBas /staining-robustness-evaluation A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models This repository provides the stain references, pretrained models, and experimental results required to: Define custom staining references using our PLISM reference library Reproduce our published controlled staining robustness experiments 👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main 👉 Associated publication: Paper Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.tabular100K<n<1M1 likes2.7k downloads4mo agoHugging Face07anonymous-stgnn-aas /TSP_EXECUTION_RUNStabular1K<n<10K1 likes2.7k downloads21d agoHugging Face08TigreGotico /arabic-stem-lexicon Arabic Diacritized-Stem Lexicon An undiacritized Arabic surface form → its most frequent diacritized stem. Standard Arabic writes no short vowels, so anything that has to pronounce Arabic must first put them back. A neural diacritizer does that well on rare words, where inference is the only thing there is. On common words it is the wrong tool: which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.tabulartext-to-speech100K<n<1M0 likes2.2k downloads2mo agoHugging Face09brandonyeequon /stock-market-data-warehousetabular100M<n<1B10 likes2.1k downloads1y agoHugging Face10snad-space /us-names-by-state US Baby names The SSA dataset with baby names: https://www.ssa.gov/OACT/babynames/ Coniferest We use this dataset in the active anomaly discovery Python package coniferest: https://coniferest.snad.space/en/latest/notebooks/us-names.html Update the data Install Python packages: pip install requests aiohttp universal_pathlib pandas Optionally: download https://www.ssa.gov/OACT/babynames/state/namesbystate.zip ./run.py PATH_OR_URL_TO_namesbystate.zip, path may be… See the full description on the dataset page: https://huggingface.co/datasets/snad-space/us-names-by-state.tabular1M<n<10M0 likes1.9k downloads1y agoHugging Face11stablellama /Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw. NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model. Raw is intended for training, as are the samples in this dataset as they can be used for regularization. Possible uses Regularization images for training models based on Krea 2 Raw Quality testing Data source This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples_Best_of.tabulartext-to-image1K<n<10K0 likes1.9k downloads22d agoHugging Face12snap-stanford /stark STaRK Website | Github | Paper STaRK is a large-scale semi-structure retrieval benchmark on Textual and Relational Knowledge Bases Downstream Task Retrieval systems driven by LLMs are tasked with extracting relevant answers from a knowledge base in response to user queries. Each knowledge base is semi-structured, featuring large-scale relational data among entities and comprehensive textual information for each entity. We have constructed three knowledge bases: Amazon SKB… See the full description on the dataset page: https://huggingface.co/datasets/snap-stanford/stark.textquestion-answering10K<n<100K12 likes1.8k downloads2y agoHugging Face13MMInstruction /stock_factorstabular10M<n<100M4 likes1.8k downloads10mo agoHugging Face14no-ry /world-stock-prices-daily-updatingtabular100K<n<1M0 likes1.8k downloads1y agoHugging Face15nadtoka /predictive-stock-datasettabular1K<n<10K0 likes1.7k downloads7h agoHugging Face16strike20023 /VideoDRtextn<1K2 likes1.5k downloads7d agoHugging Face17perctrix /StockChina-Minute A-Share Minute-Level Historical Data Dataset Description This dataset contains minute-level trading data for Chinese A-share stocks from 2005 to 2023, covering 5267 stocks with complete historical trading records. Data Format Each CSV file corresponds to one stock and contains the following fields: Field Description open Opening price close Closing price high Highest price low Lowest price volume Trading volume money Trading amount avg… See the full description on the dataset page: https://huggingface.co/datasets/perctrix/StockChina-Minute.tabular1B<n<10B14 likes1.4k downloads1y agoHugging Face18emrecan /stsb-mt-turkish STSb Turkish Semantic textual similarity dataset for the Turkish language. It is a machine translation (Azure) of the STSb English dataset. This dataset is not reviewed by expert human translators. Uploaded from this repository. Citing & Authors @misc{celik2020stsbtr, author = {Emrecan Çelik}, title = {STSB-MT-Turkish}, howpublished = {Hugging Face dataset repository}, url = {https://huggingface.co/datasets/emrecan/stsb-mt-turkish}… See the full description on the dataset page: https://huggingface.co/datasets/emrecan/stsb-mt-turkish.texttext-classification1K<n<10K8 likes1.3k downloads17d agoHugging Face19stas /gutenberg-100 Gutenberg Sci-Fi Book Dataset Testing Sample This dataset contains information about science fiction books. It’s designed for training AI models, research, or any other purpose related to natural language processing. It contains just 100 books for quick download targetting CI use (34MB). The original dataset it's derived from is https://huggingface.co/datasets/stevez80/Sci-Fi-Books-gutenberg Data Format The dataset is provided in CSV format. Each record represents a… See the full description on the dataset page: https://huggingface.co/datasets/stas/gutenberg-100.texttext-generationn<1K0 likes1.3k downloads11mo agoHugging Face20RanveerChaudhary /password_strength_datasettexttext-classification100K<n<1M2 likes1.3k downloads1y agoHugging Face21prquan /STARK_1k Benchmarking Spatiotemporal Reasoning in Large Language Models: Capabilities and Challenges Dataset for our paper: Benchmarking Spatiotemporal Reasoning in Large Language Models: Capabilities and Challenges Contact Information If you have any questions or feedback, feel free to reach out: Name: Pengrui Quan Email: prquan@ucla.edu License Copyright (c) 2025, UCLA Networked and Embedded Systems Laboratory (NESL) All rights reserved. Redistribution and use in… See the full description on the dataset page: https://huggingface.co/datasets/prquan/STARK_1k.textquestion-answering1K<n<10K0 likes1.1k downloads10mo agoHugging Face22HydraLM /glaive_function_calling_v1_standardizedtabular100K<n<1M5 likes1.1k downloads3y agoHugging Face23kawsersikder /bangladesh-stock-market-dataset Bangladesh Stock Market Dataset: 27 Years of Open-Source Dhaka Stock Exchange Data with Technical Indicators and Deep Learning Benchmarks Author: Kawser Sikder Overview A comprehensive, open-source financial dataset covering 441 publicly traded instruments across 23 industry sectors of the Dhaka Stock Exchange (DSE), Bangladesh's principal securities market. Metric Value Total Stocks 441 Total Sectors 23 Total Trading Records 1,507,388 Date Range… See the full description on the dataset page: https://huggingface.co/datasets/kawsersikder/bangladesh-stock-market-dataset.tabulartime-series-forecasting1M<n<10M1 likes1.1k downloads28d agoHugging Face24stasvinokur /cve-and-cwe-dataset-1999-2025This collection brings together every Common Vulnerabilities & Exposures (CVE) entry published in the National Vulnerability Database (NVD) from the very first identifier — CVE-1999-0001 — through all records available on 30 May 2025. It was built automatically with a Python script that calls the NVD REST API v2.0 page-by-page, handles rate-limits, and filters data. After download each CVE object is pared down to the essentials and written to CVE_CWE_2025.csv with the following columns:… See the full description on the dataset page: https://huggingface.co/datasets/stasvinokur/cve-and-cwe-dataset-1999-2025.tabulartext-classification100K<n<1M10 likes999 downloads1y agoHugging Face25xX-its-amit-Xx /pxr-structure-pose-pool PXR Structure Challenge — Full Multi-Model Pose Pool (184 ligands) Every protein–ligand pose generated during the OpenADMET PXR (pregnane X receptor / NR1I2) structure-prediction challenge, released openly with per-pose labels so the community can reuse the compute already spent — and, we hope, crack the problem this data makes visible. What's here poses/<model>/<SID>.pdb — one best pose per (model, ligand). Protein chain A + ligand (resname LIG). 15 models, up… See the full description on the dataset page: https://huggingface.co/datasets/xX-its-amit-Xx/pxr-structure-pose-pool.tabular1K<n<10K0 likes938 downloads2mo agoHugging Face26StructBench /notch-beam-2d-impact NotchBeam2D-Impact — StructBench canonical dataset Download One case, one file — fetch exactly what you need (pip install huggingface_hub): from huggingface_hub import hf_hub_download, snapshot_download # one case path = hf_hub_download("StructBench/notch-beam-2d-impact", filename="<case_id>.h5", repo_type="dataset") # the full archive (resumable; cached under HF_HOME) root = snapshot_download("StructBench/notch-beam-2d-impact"… See the full description on the dataset page: https://huggingface.co/datasets/StructBench/notch-beam-2d-impact.tabularn<1K0 likes925 downloads25d agoHugging Face27scikit-fingerprints /LRGB_Peptides-struct LRGB Peptides-struct Peptides-struct (Peptides structural) dataset, part of Long Range Graph Benchmark (LRGB) [1]. It is intended to be used through scikit-fingerprints library. The task is to predict structural properties of peptides. Note that this is raw data, whereas the original paper [1] specifies that targets should be standardized (mean 0, standard deviation 1) before training and evaluation. scikit-fingerprints does this by default in the loader function, otherwise this… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/LRGB_Peptides-struct.tabulartabular-classification10K<n<100K0 likes905 downloads6mo agoHugging Face28starmpcc /Asclepius-Synthetic-Clinical-Notes Asclepius: Synthetic Clincal Notes & Instruction Dataset Dataset Summary This dataset is official dataset for Asclepius (arxiv) This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs. We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5 Then, we generate instruction-answer pairs for 157k synthetic discharge summaries Supported Tasks This dataset covers below 8 tasks Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.textquestion-answering100K<n<1M117 likes890 downloads2y agoHugging Face29strikersoft /strikerData 🎧 StrikerData Overview StrikerData is an audio dataset developed by Strikersoft for research and development in audio and speech technologies.It contains human speech, environmental noise, and other sound types. The dataset is available for non-commercial use only, except for the company Strikersoft. Category Percentage of Total Dataset Clean human speech 20% Distorted speech 15% Human-made noise 15% Non-human noise 50% ⚖️ License… See the full description on the dataset page: https://huggingface.co/datasets/strikersoft/strikerData.audio10K<n<100K2 likes831 downloads8mo agoHugging Face30HaoranoLee /pizza_st_human_mouse_v1_train_with_labelstextn<1K0 likes822 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.