datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
c3vdv2-SfM
C3VDv2 — Colonoscopy 3D Video Dataset v2
This dataset is a re-packaged version of C3VDv2 originally published by
Johns Hopkins University, distributed under the
Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Original dataset DOI: https://doi.org/10.7281/T1/JC64MK
Dataset archive: https://archive.data.jhu.edu/dataset.xhtml?persistentId=doi:10.7281/T1/JC64MK
Attribution
This re-packaged version was created to facilitate streaming access. The… See the full description on the dataset page: https://huggingface.co/datasets/SmartWhatt/c3vdv2-SfM.smart-turn-data-v3.2-trainTraining dataset for Smart Turn v3.2.
Thank you to the following contributors whose audio samples are included in this dataset:
The Pipecat team
Liva AI: https://www.theliva.ai/
Midcentury: https://www.midcentury.xyz/
MundoAI: https://mundoai.world/
Also, thank you to the following people for the CC-0 background noise sample data which has been used in this dataset:
https://freesound.org/people/4team/sounds/214995/
https://freesound.org/people/tomhannen/sounds/698090/… See the full description on the dataset page: https://huggingface.co/datasets/pipecat-ai/smart-turn-data-v3.2-train.smartcore-v1-data
SmartCore V4 — Pretraining Verisi (12B token, EN+TR+kod+math)
SmartCore V4 (sıfırdan ~180M Mamba-3 SISO + GQA 5:1 hibrit, EN temel + TR ikincil)
projesinin ön-tokenize edilmiş, dekontamine pretraining verisi.
Tokenizer: kdirgul/smartcore-v4 → tokenizer/ (48K SentencePiece, byte_fallback, EN+TR).
Paketleme: Her doküman encode + EOS → ardışık 2048 token'lık dizilere paketlendi (doc-arası carry-over). Padding yok.
Format: parquet shard'lar, şema {input_ids: list<uint16>[2048]… See the full description on the dataset page: https://huggingface.co/datasets/kdirgul/smartcore-v1-data.Amazon_Sample_Metadata_2023
Dataset Card for Dataset Name
Original datasets can be found on: https://amazon-reviews-2023.github.io/
Dataset Details
This dataset was made as sample of several datasets from the link above.
Dataset Description
This dataset is a curated sample derived from seven filtered Amazon product category datasets(Amazon All Beauty, Amazon Fashion, Sports and Outdoors,
Health and Personal Care, Amazon Clothing Shoes and Jewlery,
Baby Products and Beauty and Personal… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/Amazon_Sample_Metadata_2023.smart-home-energy-prediction
Smart Home Appliance Energy Prediction
Dataset Summary
A public, viewer-ready educational challenge dataset. Host-only scoring data and hidden targets are excluded.
Splits
Split
Examples
Description
train
15,882
Labeled training data
test
3,853
Public inputs with withheld target labels or annotations
Data Fields
Field
Type
date
object
lights
int64
T1
float64
RH_1
float64
T2
float64
RH_2
float64… See the full description on the dataset page: https://huggingface.co/datasets/hoangbang/smart-home-energy-prediction.smart-turn-data-v3.2-testTesting dataset for Smart Turn v3.2.
Thank you to the following contributors whose audio samples are included in this dataset:
The Pipecat team
Liva AI: https://www.theliva.ai/
Midcentury: https://www.midcentury.xyz/
MundoAI: https://mundoai.world/
Also, thank you to the following people for the CC-0 background noise sample data which has been used in this dataset:
https://freesound.org/people/4team/sounds/214995/
https://freesound.org/people/tomhannen/sounds/698090/… See the full description on the dataset page: https://huggingface.co/datasets/pipecat-ai/smart-turn-data-v3.2-test.real-colon-SfM
REAL-Colon HF Triplets
This dataset is a re-packaged version of REAL-Colon for streaming
self-supervised AF-SfMLearner training. The original dataset is distributed
under CC BY 4.0.
Original dataset DOI: https://doi.org/10.25452/figshare.plus.22202866
Rows are lower-fps temporal triplets sampled from the extracted frame files that
exist on disk. If the official extraction is already subsampled, the requested
target_fps is approximated by the nearest integer step on that available… See the full description on the dataset page: https://huggingface.co/datasets/SmartWhatt/real-colon-SfM.smart-turn-data-v3.1-trainTraining dataset for Smart Turn v3.1.
Thank you to the following contributors whose audio samples are included in this dataset:
The Pipecat team
Liva AI: https://www.theliva.ai/
Midcentury: https://www.midcentury.xyz/
MundoAI: https://mundoai.world/
smart-product-pricing-2025mmlu-smart
SMART-Filtered version of MMLU dataset
This is the SMART-Filtered MMLU dataset based on methodology proposed in Improving Model Evaluation using SMART Filtering of Benchmark Datasets
The dataset is filtered using 3 main steps: removing easy examples, removing data contaminated examples and removing similar examples.
The results dataset is more efficient and captures model capabilities better than original dataset.
Citation Information… See the full description on the dataset page: https://huggingface.co/datasets/vipulgupta/mmlu-smart.smart-product-pricing-2025_test_datasmart-turn-data-v3-trainsmart-turn-data-v3.1-en-vi
Smart Turn v3.1 English/Vietnamese
Public filtered derivative of pipecat-ai/smart-turn-data-v3.1-train at d1691dd73ec827334a98b9245c47f8f2f0bac935.
Two dataset splits are published: eng and vie. Audio bytes and source
metadata are preserved. Check verification.json manually before production use.
tabula-muris-senis-bladder-smartseq2
Bladder Tissue from Tabula Muris Senis
Tabula Muris Senis is a mammalian aging single-cell gene expression dataset, downloaded from https://cellxgene.cziscience.com/collections/0b9d8a04-bb9d-44da-aa27-705bb65b54eb. This dataset represents the Bladder tissue, using the SmartSeq2 full-length mRNA library preparation method for single cells.
Code to download and process this dataset is available in: https://github.com/seanome/2025-longevity-x-ai-hackathon
Ageing is characterized by a… See the full description on the dataset page: https://huggingface.co/datasets/longevity-db/tabula-muris-senis-bladder-smartseq2.smart-retail-shelf-auditing-v1LaTeX_OCRsmart-turn-data-v3.1-testTesting dataset for Smart Turn v3.1.
Thank you to the following contributors whose audio samples are included in this dataset:
The Pipecat team
Liva AI: https://www.theliva.ai/
Midcentury: https://www.midcentury.xyz/
MundoAI: https://mundoai.world/
smart_contract_code_commentsAmazon_Sports_and_Outdoors_2023
Dataset Card for Dataset Name
Original dataset can be found on: https://amazon-reviews-2023.github.io/
Dataset Details
This dataset is downloaded from the link above, the category Sports and Outdoors meta dataset.
Dataset Description
This dataset is a refined version of the Amazon Sports and Outdoors 2023 meta dataset, which originally contained product metadata for sports and outdoors products that are sold on Amazon. The dataset includes detailed information… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/Amazon_Sports_and_Outdoors_2023.SmartDoc-QAsmart-turn-data-v3-testmlb-lineup-reactions
SmartStake MLB Lineup Reactions (2026)
Full-resolution sportsbook odds movements around every Underdog MLB lineup post of
the 2026 season, for the six player-prop markets a batting-order change moves. Every
row is one book's price for one selection at one moment, in a window around a lineup
drop. This is the raw material behind the study
"MLB Lineup Drops", a companion to
SmartStake MLB Player Prop Odds and Results.
Coverage
Events: 2,456 Underdog MLB lineup… See the full description on the dataset page: https://huggingface.co/datasets/SmartStake/mlb-lineup-reactions.Amazon_Clothing_Shoes_and_Jewelry_2023
Amazon Clothing Shoes and Jewelry 2023 Dataset
Original dataset can be found on: https://amazon-reviews-2023.github.io/
Dataset Details
This dataset is downloaded from the link above, the category Clothing Shoes and Jewelry meta dataset.
Dataset Description
The Amazon Clothing Shoes and Jewelry 2023 dataset provides information on products from a diverse ranger of categories, including main attributes like ratings, price, and desciptions.
The table below… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/Amazon_Clothing_Shoes_and_Jewelry_2023.Amazon_Sports_and_Outdoors_2018
Amazon Sports & Outdoors Dataset
Directory Structure
metadata: Contains product information.
reviews: Contains user reviews about the products.
filtered:
e5-base-v2_embeddings.jsonl: Contains "asin" and "embeddings" created with e5-base-v2.
metadata.jsonl: Contains "asin" and "text", where text is created from the title, description, brand, main category, and category.
reviews.jsonl: Contains "reviewerID", "reviewTime", and "asin". Reviews are filtered to include only… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/Amazon_Sports_and_Outdoors_2018.Amazon_Toys_and_Games_2018
Amazon Toys & Games Dataset
Directory Structure
metadata: Contains product information.
reviews: Contains user reviews about the products.
filtered:
e5-base-v2_embeddings.jsonl: Contains "asin" and "embeddings" created with e5-base-v2.
metadata.jsonl: Contains "asin" and "text", where text is created from the title, description, brand, main category, and category.
reviews.jsonl: Contains "reviewerID", "reviewTime", and "asin". Reviews are filtered to include only perfect… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/Amazon_Toys_and_Games_2018.miniloop-o-smart60k-moss16rvq
MiniLoop-O Smart Full 1M — MOSS 16-RVQ preprocessed
Source: {SOURCE_REPO}
Rows: 1,000,000 (T2A 350k / A2A 250k / I2T 400k).
Old MiniMind-O Mimi targets were decoded and re-encoded with {MOSS_CODEC_ID} into 16-RVQ MOSS codes.
Binary fields are uint16 little-endian and reshape to (frames,16).
audits-with-reasonsThis dataset builds on top of the base dataset by augmenting it using the quantized Llama3 8b instruct model by Unsloth
Namely, it:
Expands on the level of detail of the description and recommendation.
Cleans-up the code by fixing formatting and removing out-of-context comments (e.g external URLs which might confuse a model)
Adds two new fields: functionality and type (see table for more detail)
The non-vulnerable examples only have values for code, functionality and type='no vulnerability'… See the full description on the dataset page: https://huggingface.co/datasets/msc-smart-contract-auditing/audits-with-reasons.Smart_Contract_HF_Tokenized
Dataset Card for "Smart_Contract_HF_Tokenized"
More Information needed
vulnerability-severity-classificationThis dataset combines vulnerable functions (scraped from 5 auditting companies: Codehawks, ConsenSys, Cyfrin, Sherlock, Trust Security) and auddited functions with no vulnerabilities (scraped from Etherscan)
The purpose of the dataset is to enable training of classification models to discriminate between the 4 classes: none, low, medium and high.
Field
Description
1. function
Raw solidity code
2. severity
Severity of vulnerability ('none', low, medium, high)
Data… See the full description on the dataset page: https://huggingface.co/datasets/msc-smart-contract-auditing/vulnerability-severity-classification.egostation-smartphone-raw-v1
EgoStation Smartphone Raw Catalog v1
Pre-processing catalog of smartphone egocentric video recordings. Most
episodes are exposed as metadata only; a small set of representative
sample episodes are shipped with full raw video + derived data so
partners can inspect data quality before signing up for the full corpus.
Built: 2026-05-12T19:05Z
Total episodes: 5962
Total recorded time: ~693.9 hours
Format: Apache Parquet (catalog) + raw media for sample episodes
Maintained by: ZenO Labs… See the full description on the dataset page: https://huggingface.co/datasets/zeno-labs/egostation-smartphone-raw-v1.
