datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
road-issues-detection-dataset
Road Issues Detection Dataset
Dataset Summary
This comprehensive dataset contains 9,660 high-resolution RGB images categorized for road infrastructure issues detection. The dataset focuses on identifying critical urban infrastructure problems including potholes, damaged roads, broken road signs, illegal parking violations, and environmental cleanliness issues. It has been specifically organized and curated for computer vision and machine learning applications in smart… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/road-issues-detection-dataset.issuesMM-IssueLocBench
MM-IssueLoc Bench
MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization
📄 arXiv · 💻 Evaluation toolkit
Repository-level, multimodal issue → source-code localization. Given a GitHub issue (title + body + screenshots) and a target repository, retrieve the file(s) and function(s) that must be edited to resolve it — with visual evidence treated as an explicit input, decoupled from patch synthesis.
Config
Task… See the full description on the dataset page: https://huggingface.co/datasets/Jasaxion/MM-IssueLocBench.Civic_Issueslerobot-scale-rokae-3cam-60fps-datasetanony-mm-issue-loc
dummy MM-IssueLoc
eval dataset at:
File-level: data/canonical.parquet
Function-level: data/function_level.parquet
Download the full repo set:
data preview: python examples/load_dataset.py
repo source download: python scripts/download_repos.py
olmocr2-issue14-test
Document OCR using olmOCR-2-7B-1025-FP8
This dataset contains markdown-formatted OCR results from images in davanstrien/ufo-ColPali using olmOCR-2-7B.
Processing Details
Source Dataset: davanstrien/ufo-ColPali
Model: allenai/olmOCR-2-7B-1025-FP8
Number of Samples: 5
Processing Time: 0h 3m 27s
Processing Date: 2026-06-05 12:45 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split: train
Batch Size: 16
Max Model Length: 16,384… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/olmocr2-issue14-test.lerobot-tissue-rokae-3cam-60fps-datasettestlerobot_armufo-ocr-issue9-test
Document OCR using GLM-OCR
This dataset contains OCR results from images in davanstrien/ufo-ColPali using GLM-OCR, a compact 0.9B OCR model achieving SOTA performance.
Processing Details
Source Dataset: davanstrien/ufo-ColPali
Model: zai-org/GLM-OCR
Task: text recognition
Number of Samples: 5
Processing Time: 1.9 min
Processing Date: 2026-06-05 12:14 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split: train
Batch Size: 16… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/ufo-ocr-issue9-test.vllm-issue-test-set
