datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
image-for-LULU-classifier
LULU Generated Images (SD 1.4)
Synthetic images for 79 COCO object classes, generated with Stable Diffusion v1.4.
Layout (original paths)
explicit/{category}/sample_{idx:05d}_p{prompt}_s{seed}.png
implicit/{category}/sample_{idx:05d}_p{prompt}_s{seed}.png
explicit: 79 classes × 30 prompts × 100 seeds = 237,000 images (~98 GB)
implicit: validation-style prompts, fewer samples per class
Filename fields:
prompt_idx: 0–29
seed: 0–99 (explicit)… See the full description on the dataset page: https://huggingface.co/datasets/weilulobster/image-for-LULU-classifier.ai-detector-data
AI Detector Predictions Dataset
A continuously-growing collection of AI text detection predictions with optional user feedback, generated from the AI Text Detector Space.
Every time someone analyzes text or a URL on the Space, the prediction is appended to this dataset. Users can also click "Correct" or "Incorrect" to provide feedback, which gets stored alongside the prediction.
Schema
Field
Type
Description
id
string
Unique 12-char hex identifier… See the full description on the dataset page: https://huggingface.co/datasets/adaptive-classifier/ai-detector-data.ai-generated-images-classifierweed-datasetjetson1-grip-classifier-090126
jetson1-grip-classifier-090126
Recorded dataset — captured on jetson1 — 60 episodes · 4,983 frames @ 20 fps (~3 min of demonstration).
Tasks
Instruction
Episodes
grab the cucumber close to one of the cucumber's end
50
Recording
Rig
jetson1 (calibration sidecar)
Recorded
2026-09-01
Operator
dorischen
Episode sources
60 teleop
Hardware
Robot: vibeboard_follower_tilt — 7-dim action/state:… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-grip-classifier-090126.plover-classifier-qa-combined-current-ctx050
PLOVER Classifier + QA + Attribute Resolution Outputs
Each folder under runs/ is one reproducible pipeline execution. Start with the
run's README.md, then use its numbered stage folders in order.
Current organised example: runs/pilot5k_classifier_qa_synth20260808_20260811_012206/README.md
cucumber-place-classifier-eval071526-v1-trimThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "vibeboard_follower_tilt",
"total_episodes": 74,
"total_frames": 2908,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:74"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-place-classifier-eval071526-v1-trim.cucumber-place-classifier-filtered071126This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-place-classifier-filtered071126.Classifiers-Data
Text Quality Classifier Training Dataset
This dataset is specifically designed for training text quality assessment classifiers, containing annotated data from multiple high-quality corpora covering both English and Chinese texts across various professional domains.
Dataset Summary
Total Size: ~40B tokens (after sampling)
Languages: English, Chinese
Domains: General text, Mathematics, Programming, Reasoning & QA
Annotation Dimensions: Mathematical intelligence… See the full description on the dataset page: https://huggingface.co/datasets/OpenSQZ/Classifiers-Data.Genre-Classifier-Country-Per-Country
Name Dataset — Gender Classifier Parquet
Parquet conversion of philipperemy/name-dataset for first-name gender classification.
Source
Original repository: https://github.com/philipperemy/name-dataset
Original archive: name_dataset.zip
Original CSV format: first_name,last_name,gender,country_code
Converted format: first_name,gender
One Hugging Face config/subset per country code.
Cleaning
Rows are removed when:
first_name is null, empty, or… See the full description on the dataset page: https://huggingface.co/datasets/SpiceeChat/Genre-Classifier-Country-Per-Country.autotrain-data-dog-classifiers
AutoTrain Dataset for project: dog-classifiers
Dataset Descritpion
This dataset has been automatically processed by AutoTrain for project dog-classifiers.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<474x592 RGB PIL image>",
"target": 1
},
{
"image": "<474x296 RGB PIL image>",
"target": 1
}]
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/julien-c/autotrain-data-dog-classifiers.classifier_source
Dataset Card for Lapa High Quality Pretraining Dataset
Dataset Description
Dataset Summary
This dataset is a random sample of both https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality and https://huggingface.co/datasets/lapa-llm/pretraining-high-quality to transfer classifiers from English language to Ukrainian.It was used to transfer the following models from this collection https://huggingface.co/collections/lapa-llm/lapa-v012-pretraining:… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/classifier_source.tibetan-page-orientation-classifier-dataset
Tibetan Page Orientation Dataset
Covers 7 Tibetan script families: Danyig, Druma, Gyuyig, Multi-Scripts, Pedri, Tsugdri, Uchen.
Dataset composition
Each manuscript page appears twice: once as the original scan (non_flipped) and once rotated 180° (flipped). The model's task is to distinguish these two orientations.
Scripts are balanced — each of the 7 script families contributes the same number of pages (downsampled to the smallest family).
Script (script)… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-page-orientation-classifier-dataset.classifier_raw_hfharmbench_copyright_classifier_hashes
HarmBench Copyright Classifier Hashes
Original files: https://github.com/centerforaisafety/HarmBench/tree/main/data/copyright_classifier_hashes
math-classifiers-dataThis dataset is used to train kenhktsui/math-fasttext-classifier for pretraining data curation.It contains a mix of webdata and instruction response pair.Breakdown by Source:
Source
Record Count
Label
"JeanKaddour/minipile"
1000000
Others
"open-web-math/open-web-math"
306839
Math
"math-ai/StackMathQA" ("stackmathqa200k" split)
200000
Math
"open-r1/OpenR1-Math-220k"
93733
Math
"meta-math/MetaMathQA"
395000
Math
"KbsdJames/Omni-MATH"
4428
Math
harmbench_copyright_classifier_hashes
HarmBench Copyright Classifier Hashes
Original files: https://github.com/centerforaisafety/HarmBench/tree/main/data/copyright_classifier_hashes
NES-plankton-classifier-2022-dataset
NES Plankton Classifier 2022 Training Data
Image classification dataset of plankton species imaged by Imaging FlowCytobot
(IFCB) on the Northeast US Shelf. Used to train the NES Plankton Classifier 2022.
Dataset Summary
155 classes of plankton and non-plankton ROIs (including detritus, bubbles, fibers, etc.)
97,026 labeled images across train and validation splits
Images are grayscale PNGs extracted from IFCB sample files
Annotations are human-verified using the… See the full description on the dataset page: https://huggingface.co/datasets/sosiklab/NES-plankton-classifier-2022-dataset.stage2-classifier-eclipseafrican-accent-classifiereuropean-countries-classifierarxiv-classifier
arXiv Classifier Data
Usage:
from datasets import load_dataset, DownloadMode
# download from HuggingFace
dataset = load_dataset('mlcore/arxiv-classifier', name=<CONFIG NAME>)
# load from G2
dataset = load_dataset('/share/nikola/arxiv_classifier/data/arxiv-classifier', name=<CONFIG NAME>)
To force the dataset to be re-generated:
dataset = load_dataset('/share/nikola/arxiv_classifier/data/arxiv-classifier', name=<CONFIG NAME>, download_mode=DownloadMode.FORCE_REDOWNLOAD)
See:… See the full description on the dataset page: https://huggingface.co/datasets/kilian-group/arxiv-classifier.arc-agi-impabs-dpolr1e-7-beta0.01-classifiersft5e-7test-image-classifier-datasetfood-classifier-dataset
Food classifier dataset
Class counts:
.cache: 0 images
healthy: 4525 images
not_food: 1854 images
data-use-evidence-classifier-data
data-use-evidence-classifier-data
Training corpus for the evidence-tier classifier (v3): 62,054 data-use mention
spans (marked >>> mention <<< in context) with luna-judged tiers
(tier1_evidential / tier2_declaration / tier3_nonmention / junk), distilled from
rafmacalaba/data-use-mentions-v2 (val + holdout dual-pass + train-top single-pass).
train — 62,054 spans, decontaminated against the gold controls
gold — 573 human-adjudicated controls (573 gold tier labels), held out for… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/data-use-evidence-classifier-data.task1338_peixian_equity_evaluation_corpus_sentiment_classifier
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1338_peixian_equity_evaluation_corpus_sentiment_classifier
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1338_peixian_equity_evaluation_corpus_sentiment_classifier.classifier-benchmarksnoise-classifier-datasetham10000-skin-lesion-classifier-data
