datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples_Best_of.practice-radar-behavioral-health-npi-sample
New behavioral-health organization NPIs — weekly NPPES sample
A 15-row public sample from a weekly, reproducible selection of newly enumerated Type 2 behavioral-health organizations in the U.S. Centers for Medicare & Medicaid Services National Plan and Provider Enumeration System (NPPES).
Edition at a glance
Measured period: July 6–12, 2026
New Type 2 organizations screened: 2,722
Behavioral-health organizations selected: 486
States and territories represented:… See the full description on the dataset page: https://huggingface.co/datasets/unitedideas/practice-radar-behavioral-health-npi-sample.Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/SBMM75/Krea-2-Raw_samples_Best_of.Llama-slideQA-Sample-FeaturesAniGen-Sample-Dataset
AniGen Sample Data
This directory is a compact example subset of the AniGen training dataset.
What Is Included
10 examples
10 unique raw assets
Full cross-modal files for each example
A subset metadata.csv with 10 rows
The retained directory layout follows the core structure of the reference test set:
raw/
renders/
renders_cond/
skeleton/
voxels/
features/
metadata.csv
statistics.txt
latents/ (encoded by the trained slat auto-encoder)
ss_latents/ (encoded by the… See the full description on the dataset page: https://huggingface.co/datasets/VAST-AI/AniGen-Sample-Dataset.Krea-2-Raw_samplesThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
The images were created in ComfyUI with the
bf16 version
of… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples.MUG-V-Training-Samples
MUG-V Training Samples
Sample training dataset for the MUG-V 10B video generation model training framework.
Dataset Description
This dataset contains pre-processed training samples for quick-start validation and testing of the MUG-V Megatron-LM training pipeline. It includes:
VideoVAE-encoded latents (8×8×8 compressed video representations)
T5-XXL text features (4096-dim embeddings)
Training metadata CSV (sample mapping and configuration)
⚠️ Note: This is a sample… See the full description on the dataset page: https://huggingface.co/datasets/MUG-V/MUG-V-Training-Samples.point-in-time-us-equity-fundamentals-sample
Tradevo Data — honest point-in-time US equity fundamentals
Fundamentals with filed-date stamps, so a backtest only sees what was public — and restatements are flagged, not silently applied.
A free sample dataset of point-in-time US equity fundamentals, built from SEC EDGAR.
Every value is stamped with the date it first became public (first_filed), so a join that
filters by first_filed <= as_of only sees what was knowable on that date — and later
revisions are kept alongside the… See the full description on the dataset page: https://huggingface.co/datasets/Tradevodata/point-in-time-us-equity-fundamentals-sample.ambient-short-samplesreal2sim-sample-usdz-scenes
Niantic Spatial Real2Sim Sample USDZ Scenes
Sample real-world USDZ scenes and particle-field variants from Niantic Spatial for robotics simulation and physical AI workflows.
Why this exists
The real world is the best simulation.
This repository contains sample scene assets from Niantic Spatial to help developers evaluate real-to-sim robotics and physical AI workflows in NVIDIA Isaac Sim and Isaac Lab.
What's included
This repository includes two… See the full description on the dataset page: https://huggingface.co/datasets/NianticSpatial/real2sim-sample-usdz-scenes.DREAM_SAMPLE_600Kfunction_calling_v3_SAMPLE
Trelis Function Calling Dataset - VERSION 3 - SAMPLE
This is a SAMPLE of the v3 dataset available for purchase here.
Features:
Allows models to be fine-tuned for function-calling.
The dataset is human generated and does not make use of Llama 2 or OpenAI!
The dataset includes 66 training rows, 19 validation rows and 5 test rows (for manual evaluation).
Based on eight functions: search_bing, search_arxiv, save_chat, read_json_file, list_files, get_current_weather, delete_file… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/function_calling_v3_SAMPLE.sample-community-dataset
Field
Type†
What it contains
challenge_id
integer
Unique numeric identifier for the coding challenge
challenge_slug
string
URL-friendly slug used in challenge links
challenge_name
string
Human-readable challenge title
challenge_body
string
Full challenge description (HTML/Markdown) including input/output, examples, etc.
challenge_kind
string
High-level content type (e.g., code, game)
challenge_preview
string
One-sentence teaser shown in listings
challenge_category
string… See the full description on the dataset page: https://huggingface.co/datasets/hackerrank/sample-community-dataset.incidb-skincare-free-sample
INCIDB: Skincare & Cosmetics INCI Database (free sample)
Full dataset: incidb.dataengineered.io · $79 one-time (INCIDB Complete, CSV + Parquet) → Buy on Stripe · the same sample on Kaggle
This is the free sample, not the full corpus: 200 products drawn at random from a seeded eligible pool, with their brands, the 1109 ingredients they reference, all their composition links, and the matching name-map rows. Identical schema and identical columns to the paid snapshot.
A… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/incidb-skincare-free-sample.nyc_taxi_trip_2024_p1_samplepython_codes_samplerecalldb-product-recalls-sample
RecallDB — U.S. Product Recall Database (Sample)
Full dataset: recalldb.dataengineered.io · $49 one-time snapshot → Buy on Stripe · the same sample on Kaggle
127,783 official recalls · 292,790 recalled products · CPSC · FDA · FSIS · NHTSA · USCG · 100% source-linked
RecallDB is a normalized, provenance-tracked dataset of official U.S. federal product recalls. It joins five official source families into one relational model: CPSC consumer products, NHTSA vehicles, FDA/openFDA… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/recalldb-product-recalls-sample.usta-feeds-samples
Dated samples of United States public-record change files
15 samples, one folder per family. Each folder holds sample.csv, sample.json and a README naming the source, the columns, the sealing date and the row count.
Every file is a change file, not a snapshot. We seal dated copies of a public source, compare two copies, and keep what appeared, what stopped being listed, and what quietly changed in between. Most of these sources publish only the list as it stands today and… See the full description on the dataset page: https://huggingface.co/datasets/gmreincglm/usta-feeds-samples.whiskydb-fine-spirits-sample
🥃 WhiskyDB — Fine Spirits & Whisky Dataset (Free Sample)
Full dataset: whiskydb.dataengineered.io · $49 one-time (or $49 / month with the monthly refresh) → Buy once · Subscribe · the same sample on Kaggle
A free sample of WhiskyDB: a structured, relational dataset of whiskies and fine spirits built entirely from open, legally accessible public sources — government label registries (US TTB COLA), corporate registries (UK Companies House), the EU eAmbrosia GI register, Open… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/whiskydb-fine-spirits-sample.pango-sample
Pango Sample: Real-World Computer Use Agent Training Data
Pango represents Productivity Applications with Natural GUI Observations and trajectories.
Dataset Description
This dataset contains authentic computer interaction data collected from users performing real work tasks in productivity applications. The data was collected through Pango, a crowdsourced platform where users are compensated for contributing their natural computer interactions during actual work sessions.… See the full description on the dataset page: https://huggingface.co/datasets/chakra-labs/pango-sample.automotive-service-intelligence-sample
🚗 Automotive Service Intelligence Sample Dataset
Connected • Longitudinal • Feature-Engineered • Commercially Available
This repository contains a fully anonymized sample of the Growing-Moss Data Automotive Service Intelligence Dataset, a production-derived dataset built for analytics, forecasting, AI/ML, benchmarking, and commercial product development.
Unlike transactional datasets that provide isolated records, the Growing-Moss dataset delivers connected intelligence… See the full description on the dataset page: https://huggingface.co/datasets/Growing-Moss-Data/automotive-service-intelligence-sample.floradb-houseplants-care-sample
🌿 FloraDB — Houseplant Care & Pet-Toxicity Dataset (Free Sample)
Full dataset: floradb.dataengineered.io · $49 one-time → Buy on Stripe · the same sample on Kaggle
A free sample of FloraDB: a structured dataset that turns subjective houseplant care advice — "bright indirect light", "water when dry" — into quantitative engineering metrics (Lux thresholds, watering-day intervals, temperature and humidity ranges), joined to ASPCA dog/cat toxicity and grounded on the GBIF… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/floradb-houseplants-care-sample.tradedatahub-masked-preview-sample
TradeDataHub Masked Contractor Preview Sample
This is a sample/teaser dataset published by TradeDataHub, a provider of downloadable U.S. contractor datasets organized by state, trade, and city.
What this sample contains
12,952 masked preview rows corresponding to TradeDataHub's city/trade products
Columns: product_id, business, city, trade, phone_available, website_available, verification_date
Business identities are masked ("Masked business") by design: this… See the full description on the dataset page: https://huggingface.co/datasets/rsaunders/tradedatahub-masked-preview-sample.Sample-Historical-Football-Odds
⚽ Sample: World Football Historical Closing Odds
⚠️ THIS IS A 1% FREE SAMPLE DATASET ⚠️
This dataset contains a tiny fraction of our premium database to let data scientists and quants test the data structure, cleanliness, and column formatting.
💎 THE FULL MASTER DATASET (1998-2026)
If you are building professional predictive models, backtesting betting algorithms, or doing institutional-grade quantitative analysis, you need the complete picture.
The Full Master Dataset… See the full description on the dataset page: https://huggingface.co/datasets/oliviersportsdata/Sample-Historical-Football-Odds.electronics-assembly-egocentric-sample
🔌 Electronics Assembly — Egocentric Video Dataset (Sample)
This dataset is part of a larger collection of egocentric activity datasets by Verbose Tech Labs LLP. If you want the full dataset, or want access to more categories? Get in touch with us:
📞 Phone: +91 7672 000 500
💬 WhatsApp: +91 7672 000 500
📧 Email: Hello@VerboseTechLabs.com
🌐 Website: VerboseTechLabs.com
🔗 More datasets: kaggle.com/verbosetechlabsllp
Dataset Summary
First-person point-of-view… See the full description on the dataset page: https://huggingface.co/datasets/VerboseTechLabs/electronics-assembly-egocentric-sample.honeybee-samples
HoneyBee Sample Files
Sample data and resource files for the HoneyBee framework — a scalable, modular toolkit for multimodal AI in oncology.
These files are used by the HoneyBee example notebooks (clinical, pathology, radiology) and by HoneyBee's molecular processing code at runtime (Hugo_symbols.tsv is fetched on first use of DNA mutation preprocessing).
Paper: HoneyBee: A Scalable Modular Framework for Creating Multimodal Oncology Datasets with Foundational Embedding Models… See the full description on the dataset page: https://huggingface.co/datasets/Lab-Rasool/honeybee-samples.crypto-5s-market-data-adausdc-sample
ADA/USDC High-Frequency Market Microstructure Data
Free 7-Day Sample
This repository provides a free 7-day sample of a much larger privately collected high-frequency cryptocurrency market dataset.
The complete historical archive contains millions of market snapshots, with data collection starting in December 2025, across 12 crypto/USDC markets.
The public ADA/USDC sample contains:
81,579 market snapshots
97 columns
7 days of continuous historical data
20 bid + 20… See the full description on the dataset page: https://huggingface.co/datasets/rfab85/crypto-5s-market-data-adausdc-sample.WIT-es_jina-clip-v2_samplesyslog_samplecsa-clinical-stage-asset-intelligence-sample
CSA — Clinical-Stage Asset Intelligence · Free Sample
Clinical trials, FDA, and SEC — linked to the drug asset and the listed sponsor, with a
forward catalyst calendar. This is a free 150-row sample of the nearest-term
catalysts; the full snapshot carries 2,221 forward catalysts (955 linked to
124 listed sponsors) and 1,890 resolved assets.
Data, not investment advice. CSA is information, not a recommendation to buy, sell,
or hold any security. Estimated catalyst dates (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/csa-clinical-stage-asset-intelligence-sample.
