datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ner2018-NEON-beetles
Dataset Card for 2018 NEON Ethanol-preserved Ground Beetles
Collection of ethanol-preserved ground beetles (family Carabidae) collected from various NEON sites in 2018 and photographed in batches in 2022. This dataset contains both group and individual specimen images (individuals segmented from the group images). Elytra measurements of the beetle specimens (taken on the images) are also provided.
Dataset Details
This dataset is composed of a collection of 577 images… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/2018-NEON-beetles.neon4cast-scoresSnapshot of the Ecological Forecasting Initiative NEON Forecasting Challenge
Includes probabilistic forecasts, observations, and skill scores across all submitted forecasts over 5 challenge themes.
GPT_Watercolor_Anime_Style_Images
GPT Watercolor Anime Style Images
Dataset Description
This is a synthetic GPT-generated Watercolor Anime Style image dataset. It contains 120 image-caption pairs with transparent watercolor washes, soft ink linework, pale paper texture, muted colors, and traditional anime illustration scenes.
The images focus on soft watercolor washes, expressive linework, gentle lighting, quiet interiors, nature scenes, village streets, character studies, and calm storybook… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/GPT_Watercolor_Anime_Style_Images.GPT_Monet_Style_Images
GPT Monet Style Images
Dataset Description
This is a synthetic GPT-generated Monet-style image dataset. It contains 100 image-caption pairs with impressionist lighting, soft broken color, painterly atmosphere, and garden, water, street, interior, and still-life compositions.
The images focus on luminous outdoor light, loose brushwork, atmospheric color, reflective water, flowers, fields, cozy scenes, cafes, village streets, and calm impressionist subjects.… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/GPT_Monet_Style_Images.GPT_Storybook_Anime_Style_Images
GPT Storybook Anime Style Images
Dataset Description
This is a synthetic GPT-generated Storybook Anime Style image dataset. It contains 100 image-caption pairs featuring original anime-inspired characters and scenes with a warm, illustrated storybook feeling.
The images focus on expressive character moments, gentle lighting, quiet interiors, nature scenes, village streets, cozy everyday settings, and calm storybook moods. Captions commonly describe soft linework… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/GPT_Storybook_Anime_Style_Images.neon_CLBJGPT_Pointillism_Style_Images
GPT Pointillism Style Images
Dataset description
This is a synthetic GPT-generated pointillism style image dataset. It contains 70 image-caption pairs with colorful dotted texture, painterly lighting, and pointillism-inspired compositions.
The images focus on colorful stippled brushwork, luminous dotted texture, painterly lighting, cozy subjects, gardens, landscapes, interiors, and decorative still-life scenes.
Contents
images/
metadata.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/GPT_Pointillism_Style_Images.NeonTreeEvaluationmlb-player-props
SmartStake MLB Player Prop Odds and Results (2026)
Minute-by-minute MLB player prop odds from ~75 sportsbooks and exchanges over the
2026 season, with the graded outcome of each prop attached. Every row is one
book's price for one selection at one minute. This is the raw material behind
the study "Sharpest Sportsbooks for MLB Player Props".
Coverage
Odds: late March 2026 through early July 2026.
Graded outcomes: March through June (games that had settled at… See the full description on the dataset page: https://huggingface.co/datasets/Neonlightzz/mlb-player-props.smolgpt-markdown-stories
SmolGPT-Fables Stories
A deterministic, text-only corpus of 96,000 original English
Markdown stories built for SmolGPT-Fables. Every row is one complete supervised
story example with an exact prompt / completion boundary, a requested scene
count from one to six, and plain-language conditioning fields.
No model, API, browser, or network service was used to create this dataset.
Dataset summary
96,000 stories across 96,000 isolated story families
25 genres and all… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/smolgpt-markdown-stories.neo_n5ygk3rnorug65lhnb2hgl2pobsw4vdin52wo2duomwtcmjunmNeonTreeEvaluation
Preprocessed NeonTreeEvaluation Dataset
This version of the NeonTreeEvaluation dataset has been pre-processed for use in our SelvaBox paper.
Citation
If you use this dataset, please cite the original paper too:
@dataset{ben_weinstein_2022_5914554,
author = {Ben Weinstein and Sergio Marconi and Ethan White},
title = {Data for the NeonTreeEvaluation Benchmark},
month = jan,
year = 2022,
publisher = {Zenodo},
version = {0.2.2},
doi = {10.5281/zenodo.5914554}… See the full description on the dataset page: https://huggingface.co/datasets/CanopyRS/NeonTreeEvaluation.ovos-wake-word-bench-synthetic-wakewords-hey_neon
OVOS wake_word bench — synthetic-wakewords-hey_neon
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/synthetic-wakewords.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_neon.neonredact-tr
NeonRedact-TR: Turkish PII Detection Dataset
A synthetic, span-annotated dataset for detecting personally identifiable
information (PII) in Turkish text.
Open multilingual PII models now list Turkish among many languages, but they are not built for it. This dataset trains models for the parts of Turkish they miss: suffixes, Turkey-specific identifiers, and people named by role.
Models trained on it are evaluated on a separate, independently written test set: NeonRedact-TR Bench.… See the full description on the dataset page: https://huggingface.co/datasets/neondijital/neonredact-tr.orca-math-50k-shortneon_CLBJ
NEON CLBJ Data
TL;DR
This repo downloads and stores sensor data from the NSF NEON (National
Ecological Observatory Network) network for the CLBJ site (Lyndon B.
Johnson National Grassland, TX). Four scripts pull three different kinds of
NEON/NEON-adjacent data:
Script
Source
Data
scripts/download_neon_product.py
NEON /data API (neonutilities.load_by_product)
any single CSV-tabular sensor product
scripts/download_all_products.py
same, looped
all 11… See the full description on the dataset page: https://huggingface.co/datasets/johnnybwell/neon_CLBJ.amazon_reviews_multicv-tts-cleanPrebuilt-wheelseurope-who-neonatal-mortality-rate-osis000003
Neonatal mortality rate (per 1000 live births) | Europe (WHO GHO)
🇪🇺 2,201 observations · 43 Europe countries · 1954–2023 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 2,201 observations of Neonatal mortality rate (per 1000 live births) data across 43 Europe countries, spanning 1954–2023, covering 1 distinct indicators.
About the source
Source: WHO Global Health Observatory
Publisher: World Health Organization
License:… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-who-neonatal-mortality-rate-osis000003.africa-synth-maternal-health-climate-maternal-neonatal-all
Climate & Maternal-Neonatal Outcomes (SSA) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-maternal-health-climate-maternal-neonatal-all.GPT_Photoreal_3D_Anime_Style_Images
GPT Photoreal 3D Anime Style Images
Dataset Description
This is a synthetic GPT-generated Photoreal 3D Anime Style image dataset. It contains 150 image-caption pairs with square, 1024 × 1024 PNG images.
The images cover individual and group portraits, everyday activities, animals, objects, environments, and fantasy scenes. Captions describe anime-inspired forms, photorealistic surface textures, and physically based 3D lighting. Every caption begins with Photoreal… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/GPT_Photoreal_3D_Anime_Style_Images.synthetic-neonatal-birth-outcomes-vitals-WHO-0-28days
Synthetic Neonatal Birth Outcomes & Vital Signs Dataset (0-28 days) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/synthetic-neonatal-birth-outcomes-vitals-WHO-0-28days.neonfox-api-structureqald9computer-use-large
Computer Use Large
A large-scale dataset of 48,478 screen recording videos (~12,300 hours) of professional software being used, sourced from the internet. All videos have been trimmed to remove non-screen-recording content (intros, outros, talking heads, transitions) and audio has been stripped.
Dataset Summary
Category
Videos
Hours
AutoCAD
10,059
2,149
Blender
11,493
3,624
Excel
8,111
2,002
Photoshop
10,704
2,060
Salesforce
7,807
2,336
VS Code
304… See the full description on the dataset page: https://huggingface.co/datasets/NeonoV1/computer-use-large.neon-tree-crowns-dta
NEON Tree Crowns — DTA edition
A unified, species-labeled tree-crown polygon set covering 38 NEON sites,
combining algorithmic crowns from the DeepTreeAttention (DTA) pipeline with
hand-annotated bounding boxes and polygons curated by the Weecology lab.
rows
41,738 crowns
individuals
39,702 unique trees
species
234 (NEON taxonID)
sites
38 NEON sites
format
single GeoPackage (neon_crowns_dta.gpkg)
CRS
EPSG:4326 (WGS84). Native UTM zone in crs_epsg.… See the full description on the dataset page: https://huggingface.co/datasets/weecology/neon-tree-crowns-dta.language-dataset
neonatal-sepsis-care
Neonatal Sepsis & Newborn Care (Blood Culture, Pathogens, KMC, Outcomes) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/neonatal-sepsis-care.
