datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nagekinoboureiwaintaishitai
Bangumi Image Base of Nageki No Bourei Wa Intai Shitai
This is the image base of bangumi Nageki no Bourei wa Intai shitai, we detected 78 characters, 5673 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/nagekinoboureiwaintaishitai.UnCLIPImageInterpolationSamplesUnCLIPTextInterpolationSamplesTestingDataset
SciReC: Diagnostic Evaluation of Relational Reasoning in Multimodal Scientific Conversations with Adaptive Interaction
This dataset contains multimodal question-answering examples grounded in
textbook figures. Records in the figure-grounded configurations are filtered to
include only examples whose referenced image files are present in this release.
Configurations
visual: 13791 figure-grounded visual questions with resolved images.
knowledge: 13501 caption/text-grounded… See the full description on the dataset page: https://huggingface.co/datasets/Naga1289/TestingDataset.CheckpointMergerSamplesnaginoasukara
Bangumi Image Base of Nagi No Asukara
This is the image base of bangumi Nagi no Asukara, we detected 23 characters, 3162 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/naginoasukara.idiom-vision-fooling
Idioms in Misleading Visual Context
A small, densely-annotated multimodal benchmark testing whether a misleading image can push a
vision-language model toward the wrong reading of a potentially idiomatic phrase, while human
annotators stay unaffected.
Each example pairs a sentence containing a potentially idiomatic expression with an image. The
image either matches the sentence's intended reading (aligned) or depicts the opposite
reading (misleading). Annotators label how… See the full description on the dataset page: https://huggingface.co/datasets/naghamo/idiom-vision-fooling.saccade-egomotion-bench
Saccade ego-motion benchmark
The stream, the raw decision signals, and the per-frame measurements behind
Saccade — an always-on edge VLM that re-encodes only
the image patches whose change ego-motion cannot explain.
This dataset exists so the central claim can be checked without running our code.
💻 Code: https://github.com/NagaYu/saccade
🤖 Model: https://huggingface.co/NagaYu/saccade-predictor
🚀 Demo: https://huggingface.co/spaces/NagaYu/saccade
The claim, in… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/saccade-egomotion-bench.japanese-str-dataset-v1
STR Dataset
Japanese STR (Scene Text Recognition) dataset in WebDataset format.
This dataset is composed of:
Images of Japanese named entities (full names and their affiliations)
Images of sentences retrieved from Aozora Bunko (青空文庫)
and their corresponding ground truth texts.
All images are synthesized using TRDG.
Dataset Structure
Split
Samples
Shards
train
10,000,000
1000
valid
50,000
5
test
50,000
5
Total
10,100,000
1010
Usage
import… See the full description on the dataset page: https://huggingface.co/datasets/nagohachi/japanese-str-dataset-v1.HAM10000img-latex-ko
IMG2LaTeX
1. LaTeX-ko
1. 1. 설명
IM2LATEX-230k 을 LLaVA 형식에 맞게 수정한 데이터셋 입니다.
사용법은 LLaVA, KoLLaVA 참고 하시기 바랍니다.
Image-Detailed-Description-Korean
Image-Detailed-Description-Korean
LLaVA-NeXT에 적혀있는 내용중 High-Quality Knowledge Learning부분에 다음의 내용이 있습니다:
Enhanced Performance with Recaptioned Data
Models trained with recaptioned data (ReCap) datasets, show a trend of enhanced performance in tasks requiring detailed image descriptions and document understanding.
The regenerated captions, ranging from 118K to 3M, demonstrate better scaling behaviors than the original captions, consistently improve model performance across… See the full description on the dataset page: https://huggingface.co/datasets/Nagase-Kotono/Image-Detailed-Description-Korean.jawildtext_cropped
jawildtext_cropped
Per-polygon scene-text-recognition crops derived from
llm-jp/jawildtext.
Each row of the source dataset carries a polygons column with quadrilateral
text regions and their transcriptions. For every polygon we perspective-warp
the source image onto the rectified bounding rectangle, yielding a tight,
horizontally-aligned crop suitable for training/evaluating Japanese scene-text
recognition models.
Stats
Samples: 108 403
Shards: 22… See the full description on the dataset page: https://huggingface.co/datasets/nagohachi/jawildtext_cropped.NDL_pdm-ocr-part2_cropped
NDL Public Domain OCR Dataset — Line-Cropped (WebDataset)
This dataset is a derivative of the Public Domain OCR Training Dataset (FY2021) published by the National Diet Library of Japan (国立国会図書館).
Each sample is a line-level crop of a historical document page, paired with its OCR text transcription.
What was changed from the original
This dataset was not created by the National Diet Library. The following transformations were applied:
Each text line (LINE element)… See the full description on the dataset page: https://huggingface.co/datasets/nagohachi/NDL_pdm-ocr-part2_cropped.nag_datasetyu_nagabaunicontrol-testsnagato_azurlane
Dataset of nagato (Azur Lane)
This is the dataset of nagato (Azur Lane), containing 200 images and their tags.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).(LittleAppleWebUI)
Name
Images
Download
Description
raw
200
Download
Raw data with meta information.
raw-stage3
520
Download
3-stage cropped raw data with meta information.
raw-stage3-eyes
584
Download
3-stage… See the full description on the dataset page: https://huggingface.co/datasets/AppleHarem/nagato_azurlane.MidjourneySrefsFiltered subset of deepghs/midjourney_captioned_23m_full including only the prompts with the sref param passed in midjourney.
bigym-h1-native60
BiGym H1 Native Demonstrations
Canonical native H1 demonstrations for BiGym.
Branches
main: training-ready LeRobot v3, lossless embedded PNG, 50 Hz, 1-second success hold.
raw-hold3s: original 3-second LeRobot export. Archive/source only; do not use directly for new training.
Contents
Eight tasks, 60 episodes each: move_plate_h1, reach_target_single_h1, reach_target_dual_h1, reach_target_multi_modal_h1, drawer_top_close_h1, drawer_top_open_h1… See the full description on the dataset page: https://huggingface.co/datasets/Nagi-ovo/bigym-h1-native60.japanese-str-dataset-test
OCR Dataset
Japanese OCR dataset in WebDataset format.
Dataset Structure
Split
Samples
Shards
train
5,000,000
500
valid
50,000
5
test
50,000
5
Total
5,100,000
510
Usage
import webdataset as wds
base_url = "https://huggingface.co/datasets/nagohachi/japanese-str-dataset-test/resolve/main"
# Load train split
train_dataset = (
wds.WebDataset(base_url + "/train/train-{00000..00499}.tar")
.decode("pil")
.to_tuple("png", "txt")
)… See the full description on the dataset page: https://huggingface.co/datasets/nagohachi/japanese-str-dataset-test.gyurcsanynva-naganamisamamathGrabberApp_datasetindoor_window_detection_swfgrasp_bboxwindow_datasettagalog-bicol_nagaannotations_creators: []
language:
tl
bcl
language_creators:
other
license: []
multilinguality:
multilingual
pretty_name: 'Tagalog-Bicol_naga
'
size_categories:
10K<n<100K
source_datasets: []
tags: []
task_categories:
translation
task_ids: []
SpatialRGPT-Bench
