datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Castollux-Long-ParquetImages captioned using Gemini API. Some with gemini-2.0-flash-thinking-exp-1219, but most with gemini-2.0-flash-thinking-exp-01-21.
If you plan on using this for Text-to-Image training I highly suggest doing more filtering for things such as resolution and compression artifacts. I kept those images to improve Image-to-Text performance, but it would be undesirable in Text-to-Image.
DorayakiLin_parquet_move_cylinder_in_holes_mixed_v1_lerobotDorayakiLin_vggt_pick_beaker_on_mat_parquet_lerobotDanbooru-2026-parquet-metadatarefCOCOg_9k_840_sam2_parquet
refCOCOg 9k 840 with SAM2 Masks
This dataset is derived from refCOCOg_9k_840 and adds a mask column generated offline with SAM2.
Each sample contains:
id: sample identifier
problem: referring expression / query
solution: original box and point annotations
image: RGB image stored as Hugging Face image bytes
img_height: original metadata height
img_width: original metadata width
mask: SAM2-generated binary mask stored as PNG bytes
The mask column is a pseudo-label generated from… See the full description on the dataset page: https://huggingface.co/datasets/wanwan1111/refCOCOg_9k_840_sam2_parquet.example-space-to-dataset-parquetbiomedica_clinical_imagining_subset_parquet
Dataset Card for "biomedica_clinical_imagining_subset_parquet"
More Information needed
biomedica_microscopy_imagining_subset_parquet
Dataset Card for "biomedica_microscopy_imagining_subset_parquet"
More Information needed
manchu-augmented-parquet
Manchu Augmented Dataset
This dataset was exported as Parquet shards for Hugging Face Dataset Viewer compatibility.
ParquetPractice
Google/MusicCapsのデータをスペクトログラムにしたもの。
内容はmickylan2367/ColorSpectrogramと同じ(パケットファイルにしただけ)
基本的に、このリポジトリはHuggingfaceの実験場。
Traffic_Sign_Dataset_ParquetOmniDocBench-parquet
OmniDocBench Parquet
This is a Parquet-format version of the OmniDocBench v1_0 dataset.
Why This Repository Exists
The original OmniDocBench dataset contains thousands of individual image and PDF files. When downloading via huggingface_hub, each file triggers a separate HTTP request, which can lead to HuggingFace rate limiting (HTTP 429 errors) during large-scale downloads.
This Parquet version consolidates all data into a single file, reducing the number of HTTP requests… See the full description on the dataset page: https://huggingface.co/datasets/samiuc/OmniDocBench-parquet.flappy_bird_mixed_latency_parquetsynthia-rand-cityscapes-16class-parquet
SYNTHIA-RAND-CITYSCAPES 16-class Parquet
Converted from the original SYNTHIA-RAND-CITYSCAPES release.
Notes
image: RGB image bytes
label: PNG bytes of remapped segmentation mask
Label train IDs are in [0..15]
Ignore label is 255
label_format: synthia_to_cityscapes16_trainid
persian-ocr-gemini37-wins-bina-misses-parquet
Gemini 3.7 exact / Bina miss OCR crops
53 bbox crops. Columns: image and ocr.
so100_lan_v20_full_parquet_with_imagesmmc4_core_fewer_faces_parquetElliott_Bay_benthic_imagery-parquetmaps_parquet
Dataset Card for "maps_parquet"
More Information needed
tinyworlds_parquet
TinyWorlds (parquet)
Retro game frames from AlmondGod/tinyworlds,
repackaged into a uniform per-frame parquet schema with one split per game for fast,
random-access frame loading.
TinyWorlds is a Genie reimplementation: it has no action
labels and learns latent actions from video alone. These splits are therefore action-less --
ideal for training an image tokenizer (RAE/VAE) or an unconditional video world model.
Splits
split
resolution
~fps
actions… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/tinyworlds_parquet.maps_parquet_rgb
Dataset Card for "maps_parquet_rgb"
More Information needed
demon_attack_parquet2libero_parquetmcap-dishwasher-parquetbiomedica_histopathology_subset_parquetFlowGen-parquetmathcaptchaforge-dataset-21k-parquet
Dataset Card
Overview
This dataset contains labeled image crops for multi-class visual symbol classification.
Structure
images/: image files referenced by the manifests
manifest.csv: full index
train.csv, val.csv, test.csv: split manifests
Schema
All CSV files use:
id,image_path,label,width,height,split
image_path is relative to the dataset root, formatted as images/<filename>.
Usage
Load one of the split CSV files.
Resolve image_path… See the full description on the dataset page: https://huggingface.co/datasets/emrecengdev/mathcaptchaforge-dataset-21k-parquet.co3d_parquet
Probe CO3D Parquet Export
This directory contains a Parquet export of the probe-ready CO3D subset.
Summary
Source layout: experiments/probe/datasets/co3d
Export layout: experiments/probe/datasets/co3d_parquet
Row unit: one sequence with exactly 8 selected frames
Categories: 51
Selected sequences: 4396
Original on-disk sequences: 20273
Valid sequences before per-category truncation: 15938
Rows per shard: 64
Shards: train=55, val=7, test=7
Column Overview… See the full description on the dataset page: https://huggingface.co/datasets/Caesarrr/co3d_parquet.example-space-to-dataset-parquetpneuma-vision-parquetOriginal dataset: https://zenodo.org/records/7426506
ORD for the Sciences Hackathon - Vehicles Detection
[!CAUTION]
This project is an example of a hackathon project. The quality of the data produced has not been evaluated. Its goal is to provide an example on how a dataset can be update to Hugginface.
This is an example of a hackathon project presented to ORD for the sciences hackathon using the openly available pNeuma vision dataset.
Go here if you wanna know more about… See the full description on the dataset page: https://huggingface.co/datasets/katospiegel/pneuma-vision-parquet.
