datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dataset-with-standalone-yamlThis is a test dataset used in the datasets library CI
StreamingBench
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
🏠 Project Page |
📄 arXiv Paper |
📦 Dataset |
🏅Leaderboard
StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟
[NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output.
[NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/mjuicem/StreamingBench.air-bench-2024
AIRBench 2024
AIRBench 2024 is a AI safety benchmark that aligns with emerging government
regulations and company policies. It consists of diverse, malicious prompts
spanning categories of the regulation-based safety categories in the
AIR 2024 safety taxonomy.
Dataset Details
Dataset Description
AIRBench 2024 is a AI safety benchmark that aligns with emerging government
regulations and company policies. It consists of diverse, malicious prompts
spanning… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/air-bench-2024.STARK_10k
STARK: Spatial-Temporal reAsoning benchmaRK
STARK is a comprehensive benchmark designed to systematically evaluate large language models (LLMs) and large reasoning models (LRMs) on spatial-temporal reasoning tasks, particularly for applications in cyber-physical systems (CPS) such as robotics, autonomous vehicles, and smart city infrastructure.
Dataset Summary
Hierarchical Benchmark: Tasks are structured across three levels of reasoning complexity:
State Estimation:… See the full description on the dataset page: https://huggingface.co/datasets/prquan/STARK_10k.Sparkle
Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance
Ziyun Zeng, Yiqi Lin, Guoqiang Liang, and Mike Zheng Shou
📦 Dataset
Sparkle is a large-scale video background replacement dataset comprising ~140K high-quality source–edited video pairs. It is fully open-sourced at 🤗stdKonjac/Sparkle. For full methodology and dataset details, please refer to our paper.
The dataset is organized into five themes along different… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/Sparkle.staining-robustness-evaluation
A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models
This repository provides the stain references, pretrained models, and experimental results required to:
Define custom staining references using our PLISM reference library
Reproduce our published controlled staining robustness experiments
👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main
👉 Associated publication: Paper
Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.TSP_EXECUTION_RUNSarabic-stem-lexicon
Arabic Diacritized-Stem Lexicon
An undiacritized Arabic surface form → its most frequent diacritized stem.
Standard Arabic writes no short vowels, so anything that has to pronounce Arabic
must first put them back. A neural diacritizer does that well on rare words, where
inference is the only thing there is. On common words it is the wrong tool:
which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked
up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.stock-market-data-warehouseus-names-by-state
US Baby names
The SSA dataset with baby names:
https://www.ssa.gov/OACT/babynames/
Coniferest
We use this dataset in the active anomaly discovery Python package coniferest:
https://coniferest.snad.space/en/latest/notebooks/us-names.html
Update the data
Install Python packages: pip install requests aiohttp universal_pathlib pandas
Optionally: download https://www.ssa.gov/OACT/babynames/state/namesbystate.zip
./run.py PATH_OR_URL_TO_namesbystate.zip, path may be… See the full description on the dataset page: https://huggingface.co/datasets/snad-space/us-names-by-state.Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples_Best_of.stark
STaRK
Website | Github | Paper
STaRK is a large-scale semi-structure retrieval benchmark on Textual and Relational Knowledge Bases
Downstream Task
Retrieval systems driven by LLMs are tasked with extracting relevant answers from a knowledge base in response to user queries. Each knowledge base is semi-structured, featuring large-scale relational data among entities and comprehensive textual information for each entity. We have constructed three knowledge bases: Amazon SKB… See the full description on the dataset page: https://huggingface.co/datasets/snap-stanford/stark.stock_factorsworld-stock-prices-daily-updatingpredictive-stock-datasetVideoDRStockChina-Minute
A-Share Minute-Level Historical Data
Dataset Description
This dataset contains minute-level trading data for Chinese A-share stocks from 2005 to 2023, covering 5267 stocks with complete historical trading records.
Data Format
Each CSV file corresponds to one stock and contains the following fields:
Field
Description
open
Opening price
close
Closing price
high
Highest price
low
Lowest price
volume
Trading volume
money
Trading amount
avg… See the full description on the dataset page: https://huggingface.co/datasets/perctrix/StockChina-Minute.stsb-mt-turkish
STSb Turkish
Semantic textual similarity dataset for the Turkish language. It is a machine translation (Azure) of the STSb English dataset. This dataset is not reviewed by expert human translators.
Uploaded from this repository.
Citing & Authors
@misc{celik2020stsbtr,
author = {Emrecan Çelik},
title = {STSB-MT-Turkish},
howpublished = {Hugging Face dataset repository},
url = {https://huggingface.co/datasets/emrecan/stsb-mt-turkish}… See the full description on the dataset page: https://huggingface.co/datasets/emrecan/stsb-mt-turkish.gutenberg-100
Gutenberg Sci-Fi Book Dataset Testing Sample
This dataset contains information about science fiction books. It’s designed for training AI models, research, or any other purpose related to natural language processing.
It contains just 100 books for quick download targetting CI use (34MB).
The original dataset it's derived from is https://huggingface.co/datasets/stevez80/Sci-Fi-Books-gutenberg
Data Format
The dataset is provided in CSV format. Each record represents a… See the full description on the dataset page: https://huggingface.co/datasets/stas/gutenberg-100.password_strength_datasetSTARK_1k
Benchmarking Spatiotemporal Reasoning in Large Language Models: Capabilities and Challenges
Dataset for our paper: Benchmarking Spatiotemporal Reasoning in Large Language Models: Capabilities and Challenges
Contact Information
If you have any questions or feedback, feel free to reach out:
Name: Pengrui Quan
Email: prquan@ucla.edu
License
Copyright (c) 2025, UCLA Networked and Embedded Systems Laboratory (NESL)
All rights reserved.
Redistribution and use in… See the full description on the dataset page: https://huggingface.co/datasets/prquan/STARK_1k.glaive_function_calling_v1_standardizedbangladesh-stock-market-dataset
Bangladesh Stock Market Dataset: 27 Years of Open-Source Dhaka Stock Exchange Data with Technical Indicators and Deep Learning Benchmarks
Author: Kawser Sikder
Overview
A comprehensive, open-source financial dataset covering 441 publicly traded instruments across 23 industry sectors of the Dhaka Stock Exchange (DSE), Bangladesh's principal securities market.
Metric
Value
Total Stocks
441
Total Sectors
23
Total Trading Records
1,507,388
Date Range… See the full description on the dataset page: https://huggingface.co/datasets/kawsersikder/bangladesh-stock-market-dataset.cve-and-cwe-dataset-1999-2025This collection brings together every Common Vulnerabilities & Exposures (CVE) entry published in the National Vulnerability Database (NVD) from the very first identifier — CVE-1999-0001 — through all records available on 30 May 2025.
It was built automatically with a Python script that calls the NVD REST API v2.0 page-by-page, handles rate-limits, and filters data.
After download each CVE object is pared down to the essentials and written to CVE_CWE_2025.csv with the following columns:… See the full description on the dataset page: https://huggingface.co/datasets/stasvinokur/cve-and-cwe-dataset-1999-2025.pxr-structure-pose-pool
PXR Structure Challenge — Full Multi-Model Pose Pool (184 ligands)
Every protein–ligand pose generated during the OpenADMET PXR (pregnane X receptor / NR1I2)
structure-prediction challenge, released openly with per-pose labels so the community can
reuse the compute already spent — and, we hope, crack the problem this data makes visible.
What's here
poses/<model>/<SID>.pdb — one best pose per (model, ligand). Protein chain A + ligand
(resname LIG). 15 models, up… See the full description on the dataset page: https://huggingface.co/datasets/xX-its-amit-Xx/pxr-structure-pose-pool.notch-beam-2d-impact
NotchBeam2D-Impact — StructBench canonical dataset
Download
One case, one file — fetch exactly what you need (pip install huggingface_hub):
from huggingface_hub import hf_hub_download, snapshot_download
# one case
path = hf_hub_download("StructBench/notch-beam-2d-impact",
filename="<case_id>.h5", repo_type="dataset")
# the full archive (resumable; cached under HF_HOME)
root = snapshot_download("StructBench/notch-beam-2d-impact"… See the full description on the dataset page: https://huggingface.co/datasets/StructBench/notch-beam-2d-impact.LRGB_Peptides-struct
LRGB Peptides-struct
Peptides-struct (Peptides structural) dataset, part of Long Range Graph Benchmark (LRGB) [1]. It is intended to be used through
scikit-fingerprints library.
The task is to predict structural properties of peptides. Note that this is raw data, whereas the original paper [1] specifies
that targets should be standardized (mean 0, standard deviation 1) before training and evaluation. scikit-fingerprints does this by
default in the loader function, otherwise this… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/LRGB_Peptides-struct.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.strikerData
🎧 StrikerData
Overview
StrikerData is an audio dataset developed by Strikersoft for research and development in audio and speech technologies.It contains human speech, environmental noise, and other sound types. The dataset is available for non-commercial use only, except for the company Strikersoft.
Category
Percentage of Total Dataset
Clean human speech
20%
Distorted speech
15%
Human-made noise
15%
Non-human noise
50%
⚖️ License… See the full description on the dataset page: https://huggingface.co/datasets/strikersoft/strikerData.pizza_st_human_mouse_v1_train_with_labels
