datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
InternSpatialqasper-yesnosciriff-yesnoOCR-Data
OCR Text Detection and Recognition Dataset
Dataset Description
A large-scale, multi-source OCR dataset aggregating 14 public benchmarks for text detection and recognition in both scene images and handwritten documents. Each image is paired with:
Transcribed text for each text region
Bounding boxes (axis-aligned rectangles) for each text region
Polygon coordinates (precise boundary points) for each text region
The dataset is stored in HuggingFace Parquet format with… See the full description on the dataset page: https://huggingface.co/datasets/Yesianrohn/OCR-Data.Health_Benchmarks
LLM Health Benchmarks Dataset by Yesil Science
The LLM Health Benchmarks Dataset is a specialized resource for evaluating large language models (LLMs) in different medical specialties. It provides structured question-answer pairs designed to test the performance of AI models in understanding and generating domain-specific knowledge.
Primary Purpose
This dataset is built to:
Benchmark LLMs in medical specialties and subfields.
Assess the accuracy and contextual… See the full description on the dataset page: https://huggingface.co/datasets/yesilhealth/Health_Benchmarks.bioasq_7b_yesno100k_movie_reviews_from_kz
100,000+ Movie Reviews from Kazakhstan: Russian, Kazakh, and Code-Switched Texts
Dataset Summary
This repository provides a publicly available corpus of 100,502 movie reviews collected from kino.kz, spanning 2001–2025 and covering 4,943 unique movie titles. The dataset is multilingual and reflects a Kazakhstan-specific online setting where reviews are predominantly written in Russian, with smaller subsets in Kazakh and Kazakh–Russian code-switched text.
Reviews are… See the full description on the dataset page: https://huggingface.co/datasets/yeshpanovrustem/100k_movie_reviews_from_kz.nsfw1024bioasq_yesno_trainv0_n1464_test100TextMuSS-Benchmsmarco-yesnoTextMuSS-10Mpens-to-holderThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 4,
"total_frames": 2813,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:4"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yeshasdaman/pens-to-holder.ChineseHandwritingOCRdataset_v3WATER-Data
WATER-Data: Datasets for WordArt-Oriented Scene Text Recognition
WATER-Data is the official dataset release for the paper
"Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods" (ECCV 2026).
WordArt (artistic text) features highly customized fonts, textures, and layouts, making
WAordArt-oriented scene TExt Recognition (WATER) substantially more challenging
than general Scene Text Recognition (STR). The primary bottleneck for WATER is the lack of
large-scale… See the full description on the dataset page: https://huggingface.co/datasets/Yesianrohn/WATER-Data.bioasq-yesno-cleanMLT_Realstocks-YESBANK-1D-candleskaznerd
A Named Entity Recognition Dataset for Kazakh
This is a modified version of the dataset provided in the LREC 2022 paper KazNERD: Kazakh Named Entity Recognition Dataset.
The original repository for the paper can be found at https://github.com/IS2AI/KazNERD.
Tokens denoting speech disfluencies and hesitations (parenthesised) and background noise [bracketed] were removed.
A total of 2,027 duplicate sentences were removed.
Statistics for training (Train), validation (Valid)… See the full description on the dataset page: https://huggingface.co/datasets/yeshpanovrustem/kaznerd.STR-Synth
STR-Synth
Paper | GitHub
This repository serves as the supplementary dataset resource for the paper What’s Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-Evolution, dedicated to summarizing and providing the representative synthetic datasets for Scene Text Recognition (STR) used in the paper's comparative experiments.
The dataset collection in this repository aggregates 6 mainstream STR synthetic datasets: MJ, ST… See the full description on the dataset page: https://huggingface.co/datasets/Yesianrohn/STR-Synth.MLT2019task380_boolq_yes_no_question
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task380_boolq_yes_no_question
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task380_boolq_yes_no_question.reward-bench-Qwen2.5-3B-yes-noreward-bench-chatgpt-4o-latest-yes-noRIROCommonUnionST
UnionST: A Strong Synthetic Engine for Scene Text Recognition
Official data of the paper "What’s Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-Evolution".
Introduction
Scene Text Recognition (STR) relies critically on large-scale, high-quality training data. While synthetic data provides a cost-effective alternative to manually annotated real data, existing rendering-based synthetic datasets suffer… See the full description on the dataset page: https://huggingface.co/datasets/Yesianrohn/UnionST.reward-bench-Phi-3-mini-128k-instruct-yes-notask362_spolin_yesand_prompt_response_sub_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task362_spolin_yesand_prompt_response_sub_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task362_spolin_yesand_prompt_response_sub_classification.yesbro2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "koch",
"total_episodes": 5,
"total_frames": 2405,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/meteorinc/yesbro2.
