datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
low-alt-satellite-image-dataset-5k-sam3-segmented_jsonFoodSeg103
Dataset Card for FoodSeg103
Dataset Summary
FoodSeg103 is a large-scale benchmark for food image segmentation. It contains 103 food categories and 7118 images with ingredient level pixel-wise annotations. The dataset is a curated sample from Recipe1M and annotated and refined by human annotators. The dataset is split into 2 subsets: training set, validation set. The training set contains 4983 images and the validation set contains 2135 images.… See the full description on the dataset page: https://huggingface.co/datasets/json9473/FoodSeg103.funsd-json
Dataset Card for FUNSD (JSON Format)
This dataset contains the preprocessed JSON format of the FUNSD dataset, designed for form understanding in noisy scanned documents. It includes the original structure with annotations in JSON format, as per the original FUNSD dataset. This dataset is intended for document understanding tasks such as OCR, layout parsing, and key-value extraction.
Dataset Details
Dataset Description
The FUNSD (Form Understanding… See the full description on the dataset page: https://huggingface.co/datasets/Tamilnilavan/funsd-json.qwen-synth-characters-100-json-test
qwen-synth-characters-100-json-test
A 100-row bench-test slice of AbstractPhil/qwen-synth-characters
processed end-to-end by the qwen-test-runner 12-process extraction system:
the 11 deterministic specialist vision tasks (tasks_json) plus caption→JSON-schema
structuring of all three prepared captions (struct_*). Built to measure wall-clock,
schema conformance, and grounding before scaling to the full 60,847-row set.
[!IMPORTANT]
These are not real people — every image is… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters-100-json-test.glaucoma_diagnosis_json_analysissynthetic-object-relations-jsonfunsd-json
Dataset Card for FUNSD (JSON Format)
This dataset contains the preprocessed JSON format of the FUNSD dataset, designed for form understanding in noisy scanned documents. It includes the original structure with annotations in JSON format, as per the original FUNSD dataset. This dataset is intended for document understanding tasks such as OCR, layout parsing, and key-value extraction.
Dataset Details
Dataset Description
The FUNSD (Form Understanding in… See the full description on the dataset page: https://huggingface.co/datasets/davidle7/funsd-json.invoice-ocr-json
Invoice OCR Dataset
This dataset contains annotated invoice images and their corresponding OCR-extracted text in structured JSON format. The data was originally sourced from an open-source invoice dataset and processed using the GPT-4o mini model to extract relevant fields such as invoice number, date, total amount, vendor, and line items.
Dataset Details
Dataset Description
This dataset is designed to support training and evaluation of document understanding… See the full description on the dataset page: https://huggingface.co/datasets/vignesh2606/invoice-ocr-json.processed_sroie_donut_dataset_json2token
Dataset Card for "processed_sroie_donut_dataset_json2token"
More Information needed
PVT-JSON-Imageupdated-JSON-datasetreceipts-jsonStreetView-Image-Dataset-10K-json-testbwm-sft-json
Provenance
Every row carries Tülu-style lineage:
content_hash — sha256 over (system, user, assistant, site, is_noop, sha256(image bytes)); stable content identity (oversampled duplicate rows share it).
source — <collection>/<site>, collection ∈ {webarena, liveweb}, inferred per trajectory (a trajectory touching the WebArena instance host is webarena).
Decontamination (report-only; no rows removed)
Checked at revision 0d31d60542066316e57a139227875ea8607cbdb1:… See the full description on the dataset page: https://huggingface.co/datasets/photonmz/bwm-sft-json.llava_json_data_echartscharts-json-thai893_Images_Jsonformatted_sft_dataset.jsonlMIDOG_DATA_JSONimage_caption_dataset1image-to-json1low_alt_satellite_image_dataset_json_trainZiyuG_Image_Generation_CoT_dpo_jsonl_charts-json-chinesecharts-json-vietnamesecharts-json-ukrainianqwen_image_rebalance_idol_json_prompted_imagesStreetView-Image-Dataset-10K-jsonlow-alt-satellite-image-dataset-5k-sam3-segmented_json_vehiclesimage-to-json
