datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhysicalAI-SmartSpaces
Physical AI Smart Spaces Dataset
Overview
Comprehensive, annotated dataset for multi-camera tracking and 2D/3D object detection. This dataset is synthetically generated with Omniverse and Cosmos Transfer.
This dataset consists of over 280 hours of video from across nearly 1,800 cameras from indoor scenes in warehouses, hospitals, retail, and more. The dataset is time synchronized for tracking humans, forklifts, pallet trucks and Autonomous Mobile Robots (AMRs)… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-SmartSpaces.smart-contract-fiesta
Zellic 2023 Smart Contract Source Index
Zellic is making publicly available a dataset of known Ethereum mainnet smart contract source code.
Our aim is to provide a contract source code dataset that is readily available to the public to download in bulk. We believe this dataset will help advance the frontier of smart contract security research. Applications include static analysis, machine learning, and more. This effort is part of Zellic’s mission to create a world with no smart… See the full description on the dataset page: https://huggingface.co/datasets/Zellic/smart-contract-fiesta.FrontierOR
Frontier-OR Benchmark
A benchmark of 179 literature-grounded OR tasks, each packaged as a
self-contained reproducible unit: natural-language problem description,
mathematical formulation, reference Gurobi implementation, test instances,
reference solutions, and an automated feasibility checker.
Designed for evaluating LLMs on the end-to-end task of turning a research
paper's OR problem into runnable, verifiably-correct optimization code.
Dataset size note
This… See the full description on the dataset page: https://huggingface.co/datasets/SmartOR/FrontierOR.smart-bin-detect
arudaev/smart-bin-detect
Training data for Smart Bin Recognition – a validator ("is there a bin?")
and an identifier ("which bin?"). The design lives in docs/04-ml-pipeline.md
in the project repo, which is private; the manifests here carry per-image
provenance and are the authoritative record of what this dataset contains.
Every image carries provenance: source, source URL, licence, region,
capture date, annotator where known, label origin (human / machine /
legacy /… See the full description on the dataset page: https://huggingface.co/datasets/arudaev/smart-bin-detect.smart-turn-data-v3.2-trainTraining dataset for Smart Turn v3.2.
Thank you to the following contributors whose audio samples are included in this dataset:
The Pipecat team
Liva AI: https://www.theliva.ai/
Midcentury: https://www.midcentury.xyz/
MundoAI: https://mundoai.world/
Also, thank you to the following people for the CC-0 background noise sample data which has been used in this dataset:
https://freesound.org/people/4team/sounds/214995/
https://freesound.org/people/tomhannen/sounds/698090/… See the full description on the dataset page: https://huggingface.co/datasets/pipecat-ai/smart-turn-data-v3.2-train.c3vdv2-SfM
C3VDv2 — Colonoscopy 3D Video Dataset v2
This dataset is a re-packaged version of C3VDv2 originally published by
Johns Hopkins University, distributed under the
Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Original dataset DOI: https://doi.org/10.7281/T1/JC64MK
Dataset archive: https://archive.data.jhu.edu/dataset.xhtml?persistentId=doi:10.7281/T1/JC64MK
Attribution
This re-packaged version was created to facilitate streaming access. The… See the full description on the dataset page: https://huggingface.co/datasets/SmartWhatt/c3vdv2-SfM.DeepJEB-PP
DeepJEB++
Foundation Model-Driven Large-Scale 3D Engineering Dataset via 2D Latent Space Augmentation
Soyoung Yoo · Leekyo Jeong · Jinsu Ra · Dongeon Lee · Sunwoong Yang · Hyogu Jeong · Namwoo Kang — KAIST SmartDesignLab
📦 Dataset size & viewer note. DeepJEB++ contains 15,360 deployable, simulation-labeled brackets. The Hugging Face Dataset Viewer above shows only a small preview because the FEA field data are distributed as a compressed archive… See the full description on the dataset page: https://huggingface.co/datasets/KAIST-SmartDesignLab/DeepJEB-PP.guiact_smartphone_test
GUIAct Smartphone Dataset - Test Split
This is a FiftyOne dataset with 2079 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/guiact_smartphone_test")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/guiact_smartphone_test.Amazon_Sample_Metadata_2023
Dataset Card for Dataset Name
Original datasets can be found on: https://amazon-reviews-2023.github.io/
Dataset Details
This dataset was made as sample of several datasets from the link above.
Dataset Description
This dataset is a curated sample derived from seven filtered Amazon product category datasets(Amazon All Beauty, Amazon Fashion, Sports and Outdoors,
Health and Personal Care, Amazon Clothing Shoes and Jewlery,
Baby Products and Beauty and Personal… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/Amazon_Sample_Metadata_2023.smartcore-v1-data
SmartCore V4 — Pretraining Verisi (12B token, EN+TR+kod+math)
SmartCore V4 (sıfırdan ~180M Mamba-3 SISO + GQA 5:1 hibrit, EN temel + TR ikincil)
projesinin ön-tokenize edilmiş, dekontamine pretraining verisi.
Tokenizer: kdirgul/smartcore-v4 → tokenizer/ (48K SentencePiece, byte_fallback, EN+TR).
Paketleme: Her doküman encode + EOS → ardışık 2048 token'lık dizilere paketlendi (doc-arası carry-over). Padding yok.
Format: parquet shard'lar, şema {input_ids: list<uint16>[2048]… See the full description on the dataset page: https://huggingface.co/datasets/kdirgul/smartcore-v1-data.smart-home-energy-prediction
Smart Home Appliance Energy Prediction
Dataset Summary
A public, viewer-ready educational challenge dataset. Host-only scoring data and hidden targets are excluded.
Splits
Split
Examples
Description
train
15,882
Labeled training data
test
3,853
Public inputs with withheld target labels or annotations
Data Fields
Field
Type
date
object
lights
int64
T1
float64
RH_1
float64
T2
float64
RH_2
float64… See the full description on the dataset page: https://huggingface.co/datasets/hoangbang/smart-home-energy-prediction.VisionThink-Smart-Train
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
Senqiao/VisionThink-Smart-Train
This is the training dataset used for our Efficient Reasoning VLM on general VQA tasks.VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning [Paper]
Senqiao Yang,
Junyi Li,
Xin Lai,
Bei Yu,
Hengshuang Zhao,
Jiaya Jia
Highlights
Our VisionThink leverages reinforcement learning to autonomously learn whether… See the full description on the dataset page: https://huggingface.co/datasets/Senqiao/VisionThink-Smart-Train.material_fracturingSmartHarvest
SmartHarvest: Multi-Species Fruit Ripeness Detection Dataset
Dataset Description
SmartHarvest is a comprehensive multi-species fruit ripeness detection and segmentation dataset designed for precision agriculture applications. The dataset contains high-resolution images of fruits in natural garden environments with detailed polygon-based instance segmentation annotations and ripeness classifications.
Key Features
8 fruit species: Apple, cherry, cucumber… See the full description on the dataset page: https://huggingface.co/datasets/TheCoffeeAddict/SmartHarvest.smart-turn-data-v3.2-testTesting dataset for Smart Turn v3.2.
Thank you to the following contributors whose audio samples are included in this dataset:
The Pipecat team
Liva AI: https://www.theliva.ai/
Midcentury: https://www.midcentury.xyz/
MundoAI: https://mundoai.world/
Also, thank you to the following people for the CC-0 background noise sample data which has been used in this dataset:
https://freesound.org/people/4team/sounds/214995/
https://freesound.org/people/tomhannen/sounds/698090/… See the full description on the dataset page: https://huggingface.co/datasets/pipecat-ai/smart-turn-data-v3.2-test.DeepJEB
DeepJEB: 3D Deep Learning-Based Synthetic Jet Engine Bracket Dataset
This is the Hugging Face distribution of DeepJEB, a synthetic 3D jet engine
bracket dataset of 2,138 designs with paired geometry and finite-element
analysis (FEA) results, generated from the SimJEB seed set via a DeepSDF-based
generative model and an automated simulation pipeline.
This repository mirrors the official DeepJEB v1.0 release. To keep the ~68k
component files practical to host, each component… See the full description on the dataset page: https://huggingface.co/datasets/KAIST-SmartDesignLab/DeepJEB.slither-audited-smart-contractsThis dataset contains source code and deployed bytecode for Solidity Smart Contracts that have been verified on Etherscan.io, along with a classification of their vulnerabilities according to the Slither static analysis framework.mlb-player-props
SmartStake MLB Player Prop Odds and Results (2026)
Minute-by-minute MLB player prop odds from ~75 sportsbooks and exchanges over the
2026 season, with the graded outcome of each prop attached. Every row is one
book's price for one selection at one minute. This is the raw material behind
the study "Sharpest Sportsbooks for MLB Player Props".
Coverage
Odds: late March 2026 through early July 2026.
Graded outcomes: March through June (games that had settled at… See the full description on the dataset page: https://huggingface.co/datasets/SmartStake/mlb-player-props.trash-in-river-2025
Street Parade 2025 Dataset
Overview
This dataset was collected by SARA, a student initiative at ETH Zurich, to enable open research on trash presence in aquatic environments. It contains images of litter in the Limmat River in Zurich the day after the Street Parade (August 9, 2025). The dataset is intended for training and evaluating trash classification models.
Dataset summary
Collection date: August 9, 2025
Location: Kornhausbrücke, Zurich… See the full description on the dataset page: https://huggingface.co/datasets/SARA-smartphone-assisted-river-analysis/trash-in-river-2025.london_smart_meters
london_smart_meters (TsFile format)
5560 half hourly time series that represent the energy consumption readings of London households in kilowatt hour (kWh) from November 2011 to February 2014.
This repository contains the full source .tsf series from the Monash Time Series Forecasting Repository converted to Apache TsFile format.
Summary
Source dataset: Monash-University/monash_tsf
Original source: https://zenodo.org/record/4656072
Monash subset:… See the full description on the dataset page: https://huggingface.co/datasets/THULab/london_smart_meters.real-colon-SfM
REAL-Colon HF Triplets
This dataset is a re-packaged version of REAL-Colon for streaming
self-supervised AF-SfMLearner training. The original dataset is distributed
under CC BY 4.0.
Original dataset DOI: https://doi.org/10.25452/figshare.plus.22202866
Rows are lower-fps temporal triplets sampled from the extracted frame files that
exist on disk. If the official extraction is already subsampled, the requested
target_fps is approximated by the nearest integer step on that available… See the full description on the dataset page: https://huggingface.co/datasets/SmartWhatt/real-colon-SfM.smartFRACs
smartFRACs - Dataset of flow simulations in single rough fractures
A brief description of the dataset, its purpose, and what it contains.
Dataset of lattice Boltzmann and finite volume simulations for single phase laminar flow into single rough fractures.
Dataset Summary
Size: [e.g., 10,000 samples]
Languages: [e.g., English, Multilingual]
Data Type: [e.g., Text, Image, Audio, Tabular]
Use Case: [e.g., NLP, Vision, Speech Recognition]
Source: [e.g., Collected… See the full description on the dataset page: https://huggingface.co/datasets/smartFRACs/smartFRACs.smart-turn-data-v3.1-trainTraining dataset for Smart Turn v3.1.
Thank you to the following contributors whose audio samples are included in this dataset:
The Pipecat team
Liva AI: https://www.theliva.ai/
Midcentury: https://www.midcentury.xyz/
MundoAI: https://mundoai.world/
attackdex-paldeaSingle pokemon datasets containing all the attacks (from levelling or TMs) learnable by the relative monster. All the data refer to the Paldea region and they come from the project discussed in https://medium.com/@virtualmartire/i-built-an-algorithm-that-finds-the-optimal-pokemon-team-01ea152824a9.
smart-contract-audit-nonpdf-artifactssmart_contractmmlu-smart
SMART-Filtered version of MMLU dataset
This is the SMART-Filtered MMLU dataset based on methodology proposed in Improving Model Evaluation using SMART Filtering of Benchmark Datasets
The dataset is filtered using 3 main steps: removing easy examples, removing data contaminated examples and removing similar examples.
The results dataset is more efficient and captures model capabilities better than original dataset.
Citation Information… See the full description on the dataset page: https://huggingface.co/datasets/vipulgupta/mmlu-smart.smart-product-pricing-2025smart-repro-imagenet-resnet50-logitssmartdj_full
