datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
JDocQAThis unofficial dataset consists of QA pairs with images converted from the PDF files of JDocQA, a dataset focusing on chart and table understanding.
The conversion was performed using pdf2image.
The original dataset includes 1,176 examples, but 12 examples could not be converted into images. As a result, this image dataset consists of 1,164 examples in total.
We are uploading it here for use in the evaluation of llm-jp-eval-mm.
Please see the official github repo… See the full description on the dataset page: https://huggingface.co/datasets/speed/JDocQA.WAON
WAON: Large-Scale and High-Quality Japanese Image-Text Pair Dataset for Vision-Language Models
|
🤗 HuggingFace
|
📄 Paper
|
🧑💻 Code
|
Introduction
WAON is a Japanese (image, text) pair dataset containing approximately 155M examples, crawled from Common Crawl.
It is built from snapshots taken in 2025-18, 2025-08, 2024-51, 2024-42, 2024-33, and 2024-26.
The dataset is high-quality and diverse, constructed through a sophisticated… See the full description on the dataset page: https://huggingface.co/datasets/speed/WAON.relaion2B-multi-research-safe-jaThis is the subset of Japanese portion of relaion2B-multi-research-safe.
Reference
https://huggingface.co/datasets/laion/relaion2B-multi-research-safe
Bench2Drive-Speed-sampleThis is a aubset of Bench2Drive-Speed's CustomizedSpeedDataset where all clips are shown directly in folders, not .tar.gz files.
Full dataset is here.
Bench2Drive-Speed
Project Page | Paper | GitHub
Bench2Drive-Speed is a closed-loop benchmark for desired-speed conditioned autonomous driving, enabling explicit control over vehicle behavior through target speed and overtake/follow commands.
The CustomizedSpeedDataset contains 2,100 CARLA driving scenarios with expert demonstrations… See the full description on the dataset page: https://huggingface.co/datasets/Telkwevr/Bench2Drive-Speed-sample.my_first_lora_v1-datasetWAON-Bench-with-images
WAON-Bench: Japanese Cultural Image Classification Dataset
WAON-Bench is a manually curated image classification dataset designed to benchmark Vision-Language models on Japanese culture.
The dataset contains 374 classes across 8 categories: animals, buildings, events, everyday life, food, nature, scenery, and traditions. Each class includes 5 images.
How to Use
from datasets import load_dataset
ds = load_dataset("speed/WAON-Bench")
Data Collection… See the full description on the dataset page: https://huggingface.co/datasets/speed/WAON-Bench-with-images.image-speed-skating
Speed Skating Images
スピードスケートの画像データセット
Speed skating images of male skaters. 6,221 images.
Provided by Qlean Dataset (Amanaimages Inc. / 株式会社アマナイメージズ) - rights-cleared training data from Japan, released for academic research under a gated license.
Description(概要)
スピードスケートを行う男性を撮影した画像データセットです。
Dataset details(データ仕様)
Field
Value
Number of items(データ量)
6,249 files
Total size(データ容量)
41.06GB
File format(ファイル形式)
jpg
Speakers /… See the full description on the dataset page: https://huggingface.co/datasets/qleandataset/image-speed-skating.japanese-image-classification-evaluation-datasetThis dataset is https://huggingface.co/datasets/recruit-jp/japanese-image-classification-evaluation-dataset with jpg files.
7586 out of 7654 images were downloaded.
shellbench-shuffle_speed-1.0WAON-OCRwaon-cc-pair-with-punsafewaon-wiki-pair-removedData_NanoVLMshellbench-shuffle_speed-2.0shellbench-shuffle_speed-0.5imnet1k_speedboatshellbench-shuffle_speed-1.5waon-cc-pair-url-deduplicatedJoey_Zeloof_OpenML_SpeedDating
Overview:
After looking through different datasets i found one that includes data collected from a speeddating experiment that took place from 2002-2004, in which data was collected about the dates itselves, the outcome of the dates and background information of the participants.
My hope was that through the analysis of the data i would be able to improve my chances of sucsess in the dating world.
Data handling summery:
I went over the metadata of the dataset and didnt… See the full description on the dataset page: https://huggingface.co/datasets/Joe-Zeloof/Joey_Zeloof_OpenML_SpeedDating.speed-limit-signs
Speed Limit Signs Dataset
Dataset Overview
The Speed Limit Signs dataset contains labeled images of speed limit signs for use in machine learning tasks such as image classification and computer vision.It includes two splits — train and validation — to support model training and evaluation.
This dataset is publicly available on the Hugging Face Hub:https://huggingface.co/datasets/ctacke/speed-limit-signs
Dataset Structure
data/
├── images/
│ ├── train/
│… See the full description on the dataset page: https://huggingface.co/datasets/ctacke/speed-limit-signs.ericEric pics
crag-mm-single-turn-publiccrag-mm-single-turn-public-query-clswind_speed_100m_cogwind_speed_150m_cog
