datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hindi-ocr
Dataset Card for Dataset Name
Line-level text and images from https://heidata.uni-heidelberg.de/dataset.xhtml?persistentId=doi:10.11588/data/EGOKEI
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/apjanco/hindi-ocr.arocrbench_hindawiPlease see paper & code for more information:
https://github.com/mbzuai-oryx/KITAB-Bench
https://arxiv.org/abs/2502.14949
hindi_data_975_wordhinamatsuri
Bangumi Image Base of Hinamatsuri
This is the image base of bangumi Hinamatsuri, we detected 23 characters, 1820 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/hinamatsuri.IIT-Banaras-Hindu-University-Varanasi-Drone-PhotosAerial Image Capture from Drone of Banaras Hindu University as in 2017/2019
IIIT-INDIC-HW-WORDS-Hindi
IIIT-INDIC-HW-WORDS-Hindi
Dataset containing images of hand written words in Devanagari by various humans and the corresponding text of those images.
Overview
The dataset, originally developed by the Centre for Visual Information Technology (CVIT) at IIIT Hyderabad, has been transformed into Parquet format to facilitate its use in modern machine learning workflows. This dataset primarily targets recognition of handwritten Hindi words and aims to advance research and… See the full description on the dataset page: https://huggingface.co/datasets/c3rl/IIIT-INDIC-HW-WORDS-Hindi.HIndoor-8K
HIndoor-8K
HIndoor-8K is the first metrically calibrated real-world benchmark of indoor
RGB–D panoramas at native 8192×4096 (8K) resolution. It provides 49
equirectangular RGB panoramas, each paired with a sparse metric depth map
rendered from a real LiDAR point cloud, across 5 representative indoor
environments.
Released as a community resource for high-resolution 360° depth estimation.
Contents
HIndoor-8K/
├── README.md
├── ich/ # corridor
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/ushah/HIndoor-8K.hindi_data_975_lineOCR-Bench1000-Hindi
OCR-Bench1000-Hindi
1000 synthetic printed-text line images with ground-truth transcriptions,
sampled from a larger locally-held Hindi OCR training corpus.
This is a benchmark/sample release, not the full training set.
Data fields
Field
Description
file_name
relative path to the image (images/...)
text
ground-truth transcription
category
hindi_only / english_only / mixed / numeric_and_symbols
length_bucket
short / medium / long, by character count… See the full description on the dataset page: https://huggingface.co/datasets/meharuhanzz/OCR-Bench1000-Hindi.hindi-gov-vqa_beir This is a copy of https://huggingface.co/datasets/jinaai/hindi-gov-vqa reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai"… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/hindi-gov-vqa_beir.HINDI-BENGALI-MALAYALAM-ODIA-VISUAL-GENOMEhindi_vgqa_v1.0hints_of_truthclean_hin_vqamindbridge-phq9-hindi-evaluation
MindBridge Hindi PHQ-9/GAD-7 — Held-Out Evaluation (222 rows)
Held-out evaluation set used for the hierarchical kill-gate verdict:
Format ≥95% → Safety ≥90% Item-9 sensitivity → Utility ≥10pp Likert
accuracy vs base Gemma 4 E2B. Any single failure → drop fine-tune; ship
base; document honestly.
Composition
198 main evaluation rows — stratified random teacher carve-out
(persona × Likert × scale strata mirror the Phase D dad-review sample).
IN-DISTRIBUTION CAVEAT:… See the full description on the dataset page: https://huggingface.co/datasets/Huzayfah-Patel/mindbridge-phq9-hindi-evaluation.HINDI-BENGALI-MALAYALAM-ODIA-VISUAL-GENOMEfilter_hintsIndian_Traffic_VQA_Dataset🧭 Overview
Indian Traffic VQA is a real-world Visual Question Answering (VQA) dataset focusing on Indian road traffic signboards.
The dataset is designed for training and evaluating Vision-Language Models (VLMs) and VQA systems in the traffic and transportation domain.
This dataset bridges a gap between real-world Indian traffic conditions and machine understanding — ideal for research in autonomous driving, smart city AI, and traffic sign recognition under natural environments.
📦 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Hindu143/Indian_Traffic_VQA_Dataset.104320-Images-Korean-and-Hindi-OCR-Data-in-Natural-Scenes
Description
104,320장 규모의 한국어 및 힌디어 자연 환경 OCR 데이터셋입니다. 데이터는 포장지, 포스터, 티켓, 안내문, 메뉴, 건물 표지판 등 다양한 실환경에서 수집되었습니다. 다양한 자연 환경, 촬영 각도 및 조명 조건을 포함하여 데이터의 다양성을 확보했습니다.
어노테이션은 텍스트의 행 단위 폴리곤 바운딩 박스(또는 사각형/직사각형 바운딩 박스), 전사 및 텍스트 속성(언어 유형) 정보를 포함하며, 세로 방향 텍스트에 대해서도 폴리곤 바운딩 박스(또는 사각형/직사각형 바운딩 박스), 전사 및 텍스트 속성(언어 유형) 정보를 제공합니다. 본 데이터셋은 자연 환경에서의 한국어 및 힌디어 OCR 작업에 활용할 수 있습니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/ocr/1254?source=hf.kr
Data size
한국어 이미지 76,861장… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/104320-Images-Korean-and-Hindi-OCR-Data-in-Natural-Scenes.hindawi
Dataset Card for "hindawi"
More Information needed
hinh_anh
Dataset Card for "pickapic_v2"
please pay attention - the URLs will be temporariliy unavailabe - but you do not need them! we have in jpg_0 and jpg_1 the image bytes! so by downloading the dataset you already have the images!
More Information needed
Hindi-LLaVA-CC3M-Pretrain-595K-3dsssb_hindiThis dataset is curated as part of Cohere4AI project titled "Multimodal-Multilingual Exam".
Exam name: DSSSB (Delhi Subordinate Services Selection Board)
Hindi image/table MCQ questions were extracted from this website and the following PDFs:
https://drive.google.com/file/d/1rHgcfzqbbgdA3aLms2bu5jmCm_n9uBQ_/view
https://www.haryanajobs.org/wp-content/uploads/2021/12/DSSSB-Fee-Collector-Post-Code-99-20-Question-paper-With-Answer-Key-November-2021.pdf… See the full description on the dataset page: https://huggingface.co/datasets/srajwal1/dsssb_hindi.hina_aoki_real2llava_dataset_hinioai2025-onsite-concepts-hint-descriptionshindawi_fonts
Dataset Card for "hindawi_fonts"
More Information needed
chemistry-multimodal-exams-hindiUP_CET_Hindi_Multimodalhindi_VQA
Dataset Information
This dataset was filterd to be more balanced and this dataset was processed to create sentence embeddings . The embeddings were generated using a pre-trained sentence transformer model. Then, KMeans clustering was performed on the embeddings to group similar answers together. Finally, t-SNE was applied to reduce the dimensionality of the embeddings for visualization purposes. The resulting plot shows the clusters of sentence embeddings, which can be used for… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/hindi_VQA.
