datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples_Best_of.Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/SBMM75/Krea-2-Raw_samples_Best_of.Krea-2-Raw_samplesThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
The images were created in ComfyUI with the
bf16 version
of… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples.ambient-short-samplesMUG-V-Training-Samples
MUG-V Training Samples
Sample training dataset for the MUG-V 10B video generation model training framework.
Dataset Description
This dataset contains pre-processed training samples for quick-start validation and testing of the MUG-V Megatron-LM training pipeline. It includes:
VideoVAE-encoded latents (8×8×8 compressed video representations)
T5-XXL text features (4096-dim embeddings)
Training metadata CSV (sample mapping and configuration)
⚠️ Note: This is a sample… See the full description on the dataset page: https://huggingface.co/datasets/MUG-V/MUG-V-Training-Samples.usta-feeds-samples
Dated samples of United States public-record change files
15 samples, one folder per family. Each folder holds sample.csv, sample.json and a README naming the source, the columns, the sealing date and the row count.
Every file is a change file, not a snapshot. We seal dated copies of a public source, compare two copies, and keep what appeared, what stopped being listed, and what quietly changed in between. Most of these sources publish only the list as it stands today and… See the full description on the dataset page: https://huggingface.co/datasets/gmreincglm/usta-feeds-samples.honeybee-samples
HoneyBee Sample Files
Sample data and resource files for the HoneyBee framework — a scalable, modular toolkit for multimodal AI in oncology.
These files are used by the HoneyBee example notebooks (clinical, pathology, radiology) and by HoneyBee's molecular processing code at runtime (Hugo_symbols.tsv is fetched on first use of DNA mutation preprocessing).
Paper: HoneyBee: A Scalable Modular Framework for Creating Multimodal Oncology Datasets with Foundational Embedding Models… See the full description on the dataset page: https://huggingface.co/datasets/Lab-Rasool/honeybee-samples.improved-genie2-samplesus-business-locations-samples
LocationLists — US business location samples
Ten real rows from each of 30 of the 726 datasets published at locationlists.com, a catalog of 43,101,581 verified US business locations: manufacturer dealer networks, retail and restaurant chains, licensed trade contractors, healthcare providers, nonprofits and membership directories. Every dataset is compiled from the brand's own store locator or the official government registry, re-checked weekly, and sold as a flat CSV with no… See the full description on the dataset page: https://huggingface.co/datasets/locationlists/us-business-locations-samples.robotsim-retargeted-g1-samples
robotsim retargeted (G1) - 10 sample preview
10 retargeted motion CSVs from the internal robotsim v0.3 retargeted dataset
(anims_uniform_g1skel_retarget_newton_v1). Each CSV is one motion track;
rows are frames, columns are joint-space pose channels for the Unitree G1
robot. First retargeting using the Newton-based retargeter (NOVA -> G1).
Layout
CSV/<capture_date_YYMMDD>/<motion_name>__A<id>.csv
10 motions sampled randomly with seed=42 across the 125 capture dates.… See the full description on the dataset page: https://huggingface.co/datasets/ryji/robotsim-retargeted-g1-samples.supplychain-agent-input-samples
Supply Chain Agent Input Samples
Synthetic input samples for a five-agent supply-chain platform
(aizenio/supplychain-ops).
Each row is one internally consistent snapshot that satisfies the input
contract of every agent skill — 67 columns covering demand forecasting,
inventory monitoring, logistics/shipment evaluation, anomaly detection and
strategic analysis. A single row can be fed to any agent without
post-processing.
Why it exists
The agents needed realistic… See the full description on the dataset page: https://huggingface.co/datasets/Parthsoni10/supplychain-agent-input-samples.ambient-long-samplessample-structure-dataset
Sample dataset for PETAL model
This dataset is a sample dataset to test the functionalities of the PETAL model (encoder and decoder).
It is based on CASP15 dataset, see
https://predictioncenter.org/casp15/
https://github.com/Bhattacharya-Lab/CASP15
The registries folder contains the registry of CASP15 dataset (a csv file with filename, pdb_id, etc.)
banking77-representative-samples
About
This is a curated subset of 3 representative samples per class (77 classes in total) for the Banking77 dataset, as collected by a domain expert.
It was used in the paper "Making LLMs Worth Every Penny: Resource-Limited Text Classification in Banking", published in ACM ICAIF 2023 (https://arxiv.org/abs/2311.06102).
Our findings show that Few-Shot Text Classification on representative samples are better than randomly selected samples.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/helvia/banking77-representative-samples.pd_books_samplesgithub-samplesEDUMATH_model_samples
Math Word Problems from Comparison Models in EDUMATH: Generating Standards-aligned Educational Math Word Problems
This dataset contains 8,360 math word problems annotated by Gemma 3 27B IT and the EDUMATH Classifier from the models compared in EDUMATH: Generating Standards-aligned Educational Math Word Problems. Each row contains a question and answer along with the grade level and math standard(s) it was generated for and the model it was generated from. The Gemma 3 27B IT label… See the full description on the dataset page: https://huggingface.co/datasets/bryanchrist/EDUMATH_model_samples.mt_Samplesall_samples_Regressionpseudocode-decompiled-samples-smallbiglawbench_reversed_score_samplesluel-multilingual-tts-samples
Multilingual TTS Samples (Luel)
License: All Rights Reserved. Proprietary. Access only for authorized parties; no redistribution or use without permission. See LICENSE.
A multilingual text-to-speech / read-speech dataset of short scripted utterances across 7 languages. Each sample is a single-speaker recording of a written prompt, paired with rich speaker and recording metadata. Useful for TTS training and evaluation, ASR adaptation, dialect/accent studies, and read-speech… See the full description on the dataset page: https://huggingface.co/datasets/Luel-ai/luel-multilingual-tts-samples.layout-detector-flagged-samples
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/bibekyess/layout-detector-flagged-samples.tts-samples-with-artifactssample_students
Dataset Card for Holyfaith Academy Student Records
This dataset contains academic records for Class VIII and VII students. Personal Identifiable Information (PII) like phone numbers and birth years has been masked to ensure privacy.
Data Fields
Field Name
Description
CARD NO
Unique identification number for the student record.
Name
Full name of the student in uppercase.
OLD Class
Academic level (e.g., VIII or VII).
Old.UNIT
The school unit or section… See the full description on the dataset page: https://huggingface.co/datasets/Soumyajitxedu/sample_students.Norm_Malayalam_Evaluation_samples
Malayalam ASR Reference Prediction dataset
This repository contains evaluation results from the Malayalam ASR model "vrclc/Whisper_small_malayalam" using the "google/fleurs" dataset.
ASR Model Name: vrclc/Whisper_small_malayalam
Dataset: google/fleurs
Curated by: VRCLC
vrclc/Whisper_small_malayalam was trained with 50 hours of Malayalam speech data.
The test set of google/fleurs dataset which consists of Malayalam speech data was used to evaluate the model
The evaluation of 500… See the full description on the dataset page: https://huggingface.co/datasets/asr-malayalam/Norm_Malayalam_Evaluation_samples.biglawbench_samplessample_suggestion_acceptancefinetuning-sentiment-model-3000-samplesmicro-text-samples-v1
Micro Text Samples v1
A small human-written text dataset for validation and experimentation.
Intended Use
Testing, demos, and lightweight NLP experiments.
License
CC-BY-4.0
