datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
samsum
Dataset Card for SAMSum Corpus
Dataset Description
Links
Homepage: hhttps://arxiv.org/abs/1911.12237v2
Repository: https://arxiv.org/abs/1911.12237v2
Paper: https://arxiv.org/abs/1911.12237v2
Point of Contact: https://huggingface.co/knkarthick
Dataset Summary
The SAMSum dataset contains about 16k messenger-like conversations with summaries. Conversations were created and written down by linguists fluent in English. Linguists were asked to… See the full description on the dataset page: https://huggingface.co/datasets/knkarthick/samsum.practice-radar-behavioral-health-npi-sample
New behavioral-health organization NPIs — weekly NPPES sample
A 15-row public sample from a weekly, reproducible selection of newly enumerated Type 2 behavioral-health organizations in the U.S. Centers for Medicare & Medicaid Services National Plan and Provider Enumeration System (NPPES).
Edition at a glance
Measured period: July 6–12, 2026
New Type 2 organizations screened: 2,722
Behavioral-health organizations selected: 486
States and territories represented:… See the full description on the dataset page: https://huggingface.co/datasets/unitedideas/practice-radar-behavioral-health-npi-sample.Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples_Best_of.Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/SBMM75/Krea-2-Raw_samples_Best_of.Llama-slideQA-Sample-FeaturesAniGen-Sample-Dataset
AniGen Sample Data
This directory is a compact example subset of the AniGen training dataset.
What Is Included
10 examples
10 unique raw assets
Full cross-modal files for each example
A subset metadata.csv with 10 rows
The retained directory layout follows the core structure of the reference test set:
raw/
renders/
renders_cond/
skeleton/
voxels/
features/
metadata.csv
statistics.txt
latents/ (encoded by the trained slat auto-encoder)
ss_latents/ (encoded by the… See the full description on the dataset page: https://huggingface.co/datasets/VAST-AI/AniGen-Sample-Dataset.semikongbench-100
SemiKongBench-100
One hundred independently authored synthetic semiconductor operations tasks: ten workflows across ten frozen fabs. Each task includes 20 agent-visible assets, a task-local SQLite world, 38 provider-shaped tools, an oracle trajectory, before/after snapshots, and a deterministic 100-point verifier.
The model leaderboard is intentionally empty until a complete version-pinned 100-task model run exists. Release qualification executes oracle, replay, and six… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/semikongbench-100.Crop-Recommendation-Parameters
🌱 Crop Recommendation Dataset
A machine learning dataset for crop recommendation based on soil properties and environmental conditions. The dataset contains measurements of essential soil nutrients and climatic parameters, along with the crop label that is suitable for those conditions.
This dataset can be used for machine learning classification, agricultural analytics, decision-support systems, and smart farming applications.
📌 Dataset Overview
Property… See the full description on the dataset page: https://huggingface.co/datasets/Samarth-27/Crop-Recommendation-Parameters.llm-cost-same-prompt
Measured per-call LLM cost — same prompt, every model
Vendors publish prices per million tokens. Nobody publishes what one call actually costs, because
that depends on how many tokens the model chooses to emit — and on the same question models differ by
more than an order of magnitude. One model finishes a JSON extraction in 23 tokens; another writes 300.
This dataset sends a fixed set of prompts to every model at temperature 0, every night, and records
the cost computed from… See the full description on the dataset page: https://huggingface.co/datasets/mario0369/llm-cost-same-prompt.hubbench
HubBench 1.4.0
One Blobfish-authored, oracle-proven benchmark family per Harbor Hub professional-domain cluster. Every task is an employee decision worked over a dependent chain of evidence — never a lookup — against mock stateful tools over an isolated SQLite world. The agent reaches the world only through its public surfaces (MCP over streamable HTTP, a terminal tool CLI, a REST API, and a web console); a deterministic verifier (HubScore) grades the finished world from… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/hubbench.ts-satfire
Dataset Card for TS-SatFire
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
The TS-SatFire dataset is a comprehensive multi-temporal remote sensing dataset designed to cover the entire life cycle of wildfires. It provides a unified framework to support three critical and interconnected wildfire monitoring tasks: active fire detection, daily burned area… See the full description on the dataset page: https://huggingface.co/datasets/SamuelWu318/ts-satfire.Krea-2-Raw_samplesThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
The images were created in ComfyUI with the
bf16 version
of… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples.ambient-short-samplestest_many_filesMUG-V-Training-Samples
MUG-V Training Samples
Sample training dataset for the MUG-V 10B video generation model training framework.
Dataset Description
This dataset contains pre-processed training samples for quick-start validation and testing of the MUG-V Megatron-LM training pipeline. It includes:
VideoVAE-encoded latents (8×8×8 compressed video representations)
T5-XXL text features (4096-dim embeddings)
Training metadata CSV (sample mapping and configuration)
⚠️ Note: This is a sample… See the full description on the dataset page: https://huggingface.co/datasets/MUG-V/MUG-V-Training-Samples.point-in-time-us-equity-fundamentals-sample
Tradevo Data — honest point-in-time US equity fundamentals
Fundamentals with filed-date stamps, so a backtest only sees what was public — and restatements are flagged, not silently applied.
A free sample dataset of point-in-time US equity fundamentals, built from SEC EDGAR.
Every value is stamped with the date it first became public (first_filed), so a join that
filters by first_filed <= as_of only sees what was knowable on that date — and later
revisions are kept alongside the… See the full description on the dataset page: https://huggingface.co/datasets/Tradevodata/point-in-time-us-equity-fundamentals-sample.adaption-defi-wallet-risk-classification-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-defi_wallet_risk_classification
This dataset contains prompt-completion pairs for classifying the 14-day risk outcomes of DeFi wallets on various EVM networks based on behavioral features. Each sample provides wallet metrics such as transaction counts, action ratios, and concentration levels, followed by a binary risk label and a concise reasoning statement. The data is designed for… See the full description on the dataset page: https://huggingface.co/datasets/samscript18/adaption-defi-wallet-risk-classification-v1.SAMTOR_Novel_Target_Designs_GA-II
SAMTOR — SAM-competitive de novo designs (Technetium GA-II)
196 small molecules across two generations, generated de novo by the Technetium TC-43.ai engine (GA-II) and conditioned on the S-adenosylmethionine (SAM) pocket of human SAMTOR.
Each molecule was constructed against this pocket rather than selected from a compound library — docking (AutoDock Vina) came afterwards, to place and score the generated molecules in the site. With no approved drug, clinical candidate or… See the full description on the dataset page: https://huggingface.co/datasets/Tc-43/SAMTOR_Novel_Target_Designs_GA-II.real2sim-sample-usdz-scenes
Niantic Spatial Real2Sim Sample USDZ Scenes
Sample real-world USDZ scenes and particle-field variants from Niantic Spatial for robotics simulation and physical AI workflows.
Why this exists
The real world is the best simulation.
This repository contains sample scene assets from Niantic Spatial to help developers evaluate real-to-sim robotics and physical AI workflows in NVIDIA Isaac Sim and Isaac Lab.
What's included
This repository includes two… See the full description on the dataset page: https://huggingface.co/datasets/NianticSpatial/real2sim-sample-usdz-scenes.DREAM_SAMPLE_600Kfunction_calling_v3_SAMPLE
Trelis Function Calling Dataset - VERSION 3 - SAMPLE
This is a SAMPLE of the v3 dataset available for purchase here.
Features:
Allows models to be fine-tuned for function-calling.
The dataset is human generated and does not make use of Llama 2 or OpenAI!
The dataset includes 66 training rows, 19 validation rows and 5 test rows (for manual evaluation).
Based on eight functions: search_bing, search_arxiv, save_chat, read_json_file, list_files, get_current_weather, delete_file… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/function_calling_v3_SAMPLE.sample-community-dataset
Field
Type†
What it contains
challenge_id
integer
Unique numeric identifier for the coding challenge
challenge_slug
string
URL-friendly slug used in challenge links
challenge_name
string
Human-readable challenge title
challenge_body
string
Full challenge description (HTML/Markdown) including input/output, examples, etc.
challenge_kind
string
High-level content type (e.g., code, game)
challenge_preview
string
One-sentence teaser shown in listings
challenge_category
string… See the full description on the dataset page: https://huggingface.co/datasets/hackerrank/sample-community-dataset.incidb-skincare-free-sample
INCIDB: Skincare & Cosmetics INCI Database (free sample)
Full dataset: incidb.dataengineered.io · $79 one-time (INCIDB Complete, CSV + Parquet) → Buy on Stripe · the same sample on Kaggle
This is the free sample, not the full corpus: 200 products drawn at random from a seeded eligible pool, with their brands, the 1109 ingredients they reference, all their composition links, and the matching name-map rows. Identical schema and identical columns to the paid snapshot.
A… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/incidb-skincare-free-sample.nyc_taxi_trip_2024_p1_samplepython_codes_samplerecalldb-product-recalls-sample
RecallDB — U.S. Product Recall Database (Sample)
Full dataset: recalldb.dataengineered.io · $49 one-time snapshot → Buy on Stripe · the same sample on Kaggle
127,783 official recalls · 292,790 recalled products · CPSC · FDA · FSIS · NHTSA · USCG · 100% source-linked
RecallDB is a normalized, provenance-tracked dataset of official U.S. federal product recalls. It joins five official source families into one relational model: CPSC consumer products, NHTSA vehicles, FDA/openFDA… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/recalldb-product-recalls-sample.usta-feeds-samples
Dated samples of United States public-record change files
15 samples, one folder per family. Each folder holds sample.csv, sample.json and a README naming the source, the columns, the sealing date and the row count.
Every file is a change file, not a snapshot. We seal dated copies of a public source, compare two copies, and keep what appeared, what stopped being listed, and what quietly changed in between. Most of these sources publish only the list as it stands today and… See the full description on the dataset page: https://huggingface.co/datasets/gmreincglm/usta-feeds-samples.whiskydb-fine-spirits-sample
🥃 WhiskyDB — Fine Spirits & Whisky Dataset (Free Sample)
Full dataset: whiskydb.dataengineered.io · $49 one-time (or $49 / month with the monthly refresh) → Buy once · Subscribe · the same sample on Kaggle
A free sample of WhiskyDB: a structured, relational dataset of whiskies and fine spirits built entirely from open, legally accessible public sources — government label registries (US TTB COLA), corporate registries (UK Companies House), the EU eAmbrosia GI register, Open… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/whiskydb-fine-spirits-sample.stacked-samsum-1024
stacked samsum 1024
Created with the stacked-booksum repo version v0.25. It contains:
Original Dataset: copy of the base dataset
Stacked Rows: The original dataset is processed by stacking rows based on certain criteria:
Maximum Input Length: The maximum length for input sequences is 1024 tokens in the longt5 model tokenizer.
Maximum Output Length: The maximum length for output sequences is also 1024 tokens in the longt5 model tokenizer.
Special Token: The dataset utilizes the… See the full description on the dataset page: https://huggingface.co/datasets/stacked-summaries/stacked-samsum-1024.Neuro-sama-QnAThis dataset was manually created, line by line, by my tiny hand!
Why? Because I was just bored during my summer.
