datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nano-banana-pro-prompts-datasets
🖼️ Nano Banana Pro Prompt Dataset
🖼️ The ultimate Nano Banana Pro prompt dataset (6GB+). 26,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for Nano Banana Pro AI image model and the resulting generated images. The entire dataset exceeds 6GB and contains 26,000+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/nano-banana-pro-prompts-datasets.nano-omni-vlmnanopathidp-leaderboard-resultsdolma-v1_7-305B-tokenized-llama2-nanosetnano-banana
Nano-Banana Generated Images
9,457 high-quality images generated using the Nano-Banana model (Google Gemini 2.5 Flash Image Preview).
Dataset Overview
Total Images: 9,457 images
Generation Method: Nano-Banana (Google Gemini 2.5 Flash Image Preview)
Storage Format: Optimized binary (Hugging Face Image type)
File Organization: Normal large parquet files (not chunked)
License: MIT
Schema
Column
Type
Description
id
int
Unique identifier
image
Image… See the full description on the dataset page: https://huggingface.co/datasets/bitmind/nano-banana.vamos_25pct_gpt5_nano
vamos_25pct_gpt5_nano
Description
VLN Navigation dataset with 100% of tartandrive data, 50% of scand data, 25% of coda data, 100% of in-domain spot data, and 25% of annotated/augmented data using gpt5-nano. Whenever daatsets aren't 100%, they are ranked by curvature and output of length 5.
Processing Parameters
{}
Dataset Configuration
Train dataset:
mixer: mateoguaman/vlmn_tartandrive100_scand50_coda25_spot100_sub5: 1.0… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/vamos_25pct_gpt5_nano.nanochat
nanochat
nanochat is the simplest experimental harness for training LLMs. It is designed to run on a single GPU node, the code is minimal/hackable, and it covers all major LLM stages including tokenization, pretraining, finetuning, evaluation, inference, and a chat UI. For example, you can train your own GPT-2 capability LLM (which cost $43,000 to train in 2019) for only $48 (2 hours of 8XH100 GPU node) and then talk to it in a familiar ChatGPT-like web UI. On a spot instance… See the full description on the dataset page: https://huggingface.co/datasets/ArchaeonSeq/nanochat.gpt-5-nanoNanoVDR-Train
NanoVDR-Train: Multilingual Visual Document Retrieval Training Data
Training dataset for NanoVDR, comprising 1.49M query–image pairs across 6 languages for visual document retrieval.
Paper: Our arxiv preprint is currently on hold. Details on training methodology, ablations, and full results will be available once the paper is published.
Dataset Summary
Statistic
Value
Total samples
1,489,252 (711K original + 778K augmented)
Validation samples
14,518… See the full description on the dataset page: https://huggingface.co/datasets/nanovdr/NanoVDR-Train.FFHQ-2048
FFHQ-2048 — NanoPocket Enhanced (First 1,000)
The first new high-quality public face dataset since 2019.
1,000 sharp, artifact-free 2048×2048 portraits, derived from FFHQ and enhanced with the NanoPocket Face Enhance model.
Why this dataset exists
FFHQ (NVIDIA, 2019) has been the gold-standard face dataset for the past five years — but the field has moved on. Modern generators (Flux, SD3 / SDXL, StyleGAN-T, portrait restoration nets) train at 1024²… See the full description on the dataset page: https://huggingface.co/datasets/Nanopocket-ai/FFHQ-2048.imagesnanog-cancer-data
NanoG - Cancer Foundation-Model Training Data
Multimodal cancer corpus for NanoG1 (generative multimodal pretraining). Literature, structured biology, imaging, and grounded <simulate> traces.
Hub: Abd0r/nanog-cancer-dataAuthor: Syed Abdur Rehman Ali (@Abd0r) · 17 · independent
Train exclusion (hard): NCI-60 is out of training. Skip records whose source / path / text refer to NCI-60. Prefer NCI-ALMANAC, TCGA-sim, Polymathic, PMC/PubMed, TCGA omics, imaging.
How… See the full description on the dataset page: https://huggingface.co/datasets/Abd0r/nanog-cancer-data.nanopath-evals
NanoPath evaluation data
This is the immutable data mirror used by NanoPath probe protocol v2. It contains only the exact development records consumed by medarc/nanopath: selected THUNDER training/validation images, prepared development-only slide caches, and the two PathoROB subsets. manifest.json records SHA-256 checksums and binds the snapshot to the checked-in benchmark manifests.
No official THUNDER, HEST, or CPTAC classification test record is included. HEST is absent.… See the full description on the dataset page: https://huggingface.co/datasets/medarc/nanopath-evals.nano-banana-pro-generated-1k
Nano Banana Pro (1K) Dataset
200 AI-generated images at 1K quality.
License: MIT
ade20k-nanoNano-banana-150kNano-consistent-150k. — the first dataset constructed using Nano-Banana that exceeds 150k high-quality samples, uniquely designed to preserve consistent human identity across diverse and complex editing scenarios
midjourney-dalle-sd-nanobananapro-dataset
Dataset Card: Midjourney, DALL-E, Stable Diffusion & Nano Banana Pro vs Real Images
Description
Dataset de classification binaire pour détecter les images générées par IA (Midjourney, DALL-E, Stable Diffusion et Nano Banana Pro) vs images réelles.
Dataset Structure
Train set: 10,000 images
Real: 5000 images
Fake (AI-generated): 5000 images
Test set: 2,000 images
Real: 1000 images
Fake (AI-generated): 1000 images
Features
{
"image": Image… See the full description on the dataset page: https://huggingface.co/datasets/julienlucas/midjourney-dalle-sd-nanobananapro-dataset.monuments_nanonano-receipts
🧾 Nano Receipts Dataset
A diverse collection of 2428 hyper-realistic synthetic receipt images generated using state-of-the-art text-to-image AI models.
🚀 Quick Start
from datasets import load_dataset
# Load dataset (fast parquet format!)
dataset = load_dataset("34data/nano-receipts")
# Access images
image = dataset["train"][0]["image"] # PIL Image
filename = dataset["train"][0]["filename"]
📊 Dataset Details
Total Images: 2428 receipts
Format:… See the full description on the dataset page: https://huggingface.co/datasets/34data/nano-receipts.nano-banana-pro-generated-1k-clonename: nano-banana-pro-generated-1k-clone
license: mit
pipeline_tag: text-to-image
tasks:
- text-to-image
- image-generation
tags:
- nano-banana
language: en
size_categories:
- n<1K
[!IMPORTANT]
Original Model Link : https://huggingface.co/datasets/ash12321/nano-banana-pro-generated-1k-clonenano-banana-pro-gen-zh-enFlameF0X/nano-banana-pro-gen-zh-en is a translation of kaupane/nano-banana-pro-gen from Chinese to English
nano-banana-pro-genmagical-girl-lyrical-nanoha-official-art-vernano-receipts
🧾 Nano Receipts Dataset
A diverse collection of 2428 hyper-realistic synthetic receipt images generated using state-of-the-art text-to-image AI models.
🚀 Quick Start
from datasets import load_dataset
# Load dataset (fast parquet format!)
dataset = load_dataset("34data/nano-receipts")
# Access images
image = dataset["train"][0]["image"] # PIL Image
filename = dataset["train"][0]["filename"]
📊 Dataset Details
Total Images: 2428 receipts… See the full description on the dataset page: https://huggingface.co/datasets/samarth010/nano-receipts.nanostep-datasetsnano_resultnanoNanoSRGUI-Net-NanoFor debugging purpose of training TongUI
