datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
webdsh-images
webdsh-images
Disk images for the emulated machines webdsh
offers.
Why this exists
v86 can run about a hundred and twenty-five
machines in a browser, and every one of them is the same emulator with a
different disk. What copy.sh/v86 has that a fork does
not is a CDN with the disks on it: its own host, i.copy.sh, refuses browser
requests from anywhere else — deliberately, and it is their bandwidth to
protect.
So webdsh's catalog was complete and its machines were… See the full description on the dataset page: https://huggingface.co/datasets/AndyZijianZhang/webdsh-images.rule34lol-images-part2
Dataset Card for rule34lol-images-part2
Dataset Summary
This dataset contains information about image files from rule34.lol, a booru-style imageboard. The dataset includes metadata for 77,000 image files, including URLs, tags, file information, and like counts. The actual image files are stored in zip archives, with each archive containing 1000 image files (except the last archive). This is Part 2 of 2 for the complete rule34lol-images dataset. Part 1 can be found here.… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/rule34lol-images-part2.LoRA-Merge-Imagesmed-imagesarxiv-chandra-ocr-2-include-images-first50-20260415
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415
Source paper IDs in input list: 27,584
Processed IDs recorded in state/processed_ids.txt: 50
Successes: 50
Partial successes: 0
Errors: 0
Next shard index: 10
Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415.bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bghira/pseudo-camera-10k with responses/captions generated with gemini-2.0-flash-thinking-exp-1219.
The format should be similar to that of liuhaotian/LLaVA-Instruct-150K.
Images can be found in the images.zip folder. The zip also contains .txt captions for ease of use in non-VQA tasks.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Images/bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.android-17-wasm-imagesclker-images
Dataset Card for Clker.com Images
Dataset Summary
This dataset contains 140,313 public domain clipart images collected from Clker.com. Clker.com hosts user-shared vector clip art that is explicitly released into the public domain (CC0). The dataset includes the images themselves along with metadata such as titles and tags associated with each image.
Languages
The dataset is primarily monolingual:
English (en): All image titles and tags are in… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/clker-images.rule34lol-images-part1
Dataset Card for rule34lol-images-part1
Dataset Summary
This dataset contains information about image files from rule34.lol, a booru-style imageboard. The dataset includes metadata for 196,000 image files, including URLs, tags, file information, and like counts. The actual image files are stored in zip archives, with each archive containing 1000 image files. This is Part 1 of 2 for the complete rule34lol-images dataset. Part 2 can be found here.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/rule34lol-images-part1.high_quality_images_embeddingsreal-vehicle-images-with-visual-metadata-100k
真实车辆图片与视觉元信息测试数据集
数据集简介
本数据集可扩展至 10 万张真实车辆图片,本次上传人工筛选后的 1,000 张测试数据,覆盖轿车、公交车、客车和卡车,可用于车辆图像分类、目标检测、视觉识别、多模态训练及相关算法测试。
当前版本为测试数据集。如需更多图片、指定车辆类别或定制元信息,可扩展量级,欢迎通过邮件联系:xiehan@mykj.club。
数据概况
图片数量:1,000 张
图片格式:JPEG
最小短边:513px,全部不低于 512px
类别分布:109 张轿车、442 张公交车或客车、449 张卡车
文件命名:每张图片使用唯一文件名
内容去重:使用 SHA-256 进行内容级去重
标注形式:每张图片对应一个同名 JSON 文件
目录结构
vehicle_dataset_1000/
├── README.md
├── images/ # 车辆图片
│ └── vehicle_xxx.jpg
├──… See the full description on the dataset page: https://huggingface.co/datasets/MYtechnology/real-vehicle-images-with-visual-metadata-100k.Handpicked-Images-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
Handpicked-Images-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
Some random images with responses/captions generated with gemini-2.0-flash-thinking-exp-1219.
The format should be similar to that of liuhaotian/LLaVA-Instruct-150K.
Images can be found in the images.zip folder. The zip also contains .txt captions for ease of use in non-VQA tasks.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Images/Handpicked-Images-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.civitai_top_10000_imagesimage-setr_abandoned-gemini-2.0-flash-thinking-exp-1219-CustomShareGPTarxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard index: 1
Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416.r_portraitphotography-gemini-2.0-flash-thinking-exp-1219-CustomShareGPTarxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard index: 1
Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416.arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard index: 1
Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417.fictional_characters_raw_data_with_imagesOCR-Imagesr_analog-gemini-2.0-flash-thinking-exp-1219-CustomShareGPTadaption-crop-disease-leaf-images
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-crop_disease_leaf_images
This dataset contains image-based samples for identifying various crop diseases affecting plants such as apples, corn, tomatoes, and rice. Each entry consists of a prompt requesting disease identification and a completion specifying the diagnosed condition, including healthy states. The data is formatted as prompt-completion pairs suitable for… See the full description on the dataset page: https://huggingface.co/datasets/RatnambarBaghel/adaption-crop-disease-leaf-images.arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416.adaption-crop-disease-leaf-images-v1
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-crop_disease_leaf_images
This dataset contains image-based samples for identifying various crop diseases affecting plants such as apples, corn, tomatoes, and rice. Each entry consists of a prompt requesting disease identification and a completion specifying the diagnosed condition, including healthy states. The data is formatted as prompt-completion pairs suitable for… See the full description on the dataset page: https://huggingface.co/datasets/RatnambarBaghel/adaption-crop-disease-leaf-images-v1.catalogue_snapshot_imagespixmo_images_badtransit-guest-images
Transit guest images
Prebuilt guest artifacts for Transit, downloaded on demand at VM creation time.
catalog.json and its detached catalog.sig are the machine-readable, Ed25519-signed release catalog.
Transit verifies the exact catalog bytes before parsing it and downloads every artifact from an immutable Hugging Face commit URL.
Artifacts live under <preset-id>/<fileName>.xz; the signed catalog records compressed and decompressed SHA256 digests.
Current presets:… See the full description on the dataset page: https://huggingface.co/datasets/YoozLabs/transit-guest-images.imagesarxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard index: 1… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416.
