datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OnceUponATime-florence2-captionssoa-full-florence2
Smithsonian Open Access Dataset with Florence-2 Caption
日本語はこちら
This dataset is made of soa-full.
soa-full is an CC-0 image dataset from Smithsonian Open Access. However, the dataset does not contain the image caption.
Therefore, we caption the images by Florence 2.
Usage
from datasets import load_dataset
dataset = load_dataset("aipicasso/soa-full-florence2")
Intended Use
Research Vision & Language
Develop text-to-image model or image-to-text model.… See the full description on the dataset page: https://huggingface.co/datasets/aipicasso/soa-full-florence2.danbooru2023-florence2-caption
Danbooru2023 - Florence2 Caption dataset
This dataset contains captions of danbooru2023 images generated by microsoft/Florence-2-large
I use original one with task token
Format
parquet:
key: the danbooru id of the image
parsed: parsed florence 2 output of the image
Stat
MORE_DETAILED_CAPTION
Entries: 7,438,449
Output Tokens (Min/Max/Mean/Median):
Flan T5 Tokenizer: 19/736/120/114
DFN CLIP Tokenizer: 19/826/108.7/103
Qwen2 Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/KBlueLeaf/danbooru2023-florence2-caption.megalith-10m-florence2
Megalith-10M with Florence-2 Caption
日本語はこちら
This reposity is the supplymentary of Megalith-10M.
Megalith-10M is an CC-0 like image dataset. However, the dataset does not contain the image caption.
Therefore, we caption the images by Florence 2.
Usage
from datasets import load_dataset
dataset = load_dataset("aipicasso/megalith-10m-florence2")
How to get images
git lfs install
git clone https://huggingface.co/datasets/drawthingsai/megalith-10m… See the full description on the dataset page: https://huggingface.co/datasets/aipicasso/megalith-10m-florence2.roboflow100-bccd-florence2
Dataset Card for roboflow-bccd-florrence2
Dataset Summary
This dataset, roboflow-bccd-paligemma, is a modified version of the BCCD (Blood Cell Count and Detection) dataset. It contains blood cell images annotated for object detection tasks, specifically targeting three types of blood cells:
Platelets
Red Blood Cells (RBC)
White Blood Cells (WBC)
Key features of the dataset:
Total of 364 annotated images across train, validation, and test splits
Bounding box annotations… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/roboflow100-bccd-florence2.florence2-ofa-captions-500
OFA Florence-2 Dataset (500 Samples)
This dataset was generated using microsoft/Florence-2-large on a subset of COCO 2017 Validation images.
It is pre-formatted for OFA Stage-1 fine-tuning (Headerless TSV, URL-safe base64, max 512x512 resolution).
florence2_invoice_datasetparasite_egg_dataset_florence2_detectionFFB_Florence2parasitic_egg_florence2roboflow100-bccd-florence2-tempflorence-2-minimalissimoflorence2florence-2-train-29janflorence-2-train
