datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
masscount-cf
MassCount-CF
Synthetic counterfactual counting corpus for identifying an additive
neural-mass cardinality coordinate in vision-language models (NMCA).
Companion dataset to "Counting Requires Mass: An Algebraic and Causal Account
of Numerosity in Vision-Language Models."
Corpus version: masscount-cf-1.0.0
Master scenes: 99,989
Delivered images (this upload): 326,623
Object instances: 28,617,018
Structure
Scene graphs are the durable artifact; pixels are regenerable… See the full description on the dataset page: https://huggingface.co/datasets/Ritabrata04/masscount-cf.icub_sim_dataset_t2_smooth_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "icub",
"total_episodes": 90,
"total_frames": 9935,
"total_tasks": 5,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ritamota/icub_sim_dataset_t2_smooth_lerobot.icub_sim_dataset_t4_smooth_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "icub",
"total_episodes": 90,
"total_frames": 9216,
"total_tasks": 90,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ritamota/icub_sim_dataset_t4_smooth_lerobot.icub_sim_dataset_t3_smooth_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "icub",
"total_episodes": 90,
"total_frames": 4000,
"total_tasks": 90,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ritamota/icub_sim_dataset_t3_smooth_lerobot.icub_sim_dataset_t1_smooth_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "icub",
"total_episodes": 90,
"total_frames": 12632,
"total_tasks": 64,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ritamota/icub_sim_dataset_t1_smooth_lerobot.icub_sim_dataset_t5_smooth_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "icub",
"total_episodes": 90,
"total_frames": 10716,
"total_tasks": 90,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ritamota/icub_sim_dataset_t5_smooth_lerobot.counterfactual_dataset_20_classes_x_100_samplesicub_sim_dataset_t3_simple_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "icub",
"total_episodes": 90,
"total_frames": 19830,
"total_tasks": 5,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ritamota/icub_sim_dataset_t3_simple_lerobot.mmrag-eval
mmrag-eval
Benchmark dataset for evaluating grounding quality in multimodal Retrieval-Augmented Generation (RAG) systems.
Standard benchmarks measure whether a RAG system retrieves the right document. mmrag-eval measures whether the system's generated answer faithfully reflects what is actually shown in the retrieved image — catching hallucination and retrieval redundancy that retrieval metrics alone cannot detect.
Dataset Summary
198 annotated image–query pairs… See the full description on the dataset page: https://huggingface.co/datasets/ritaban-b/mmrag-eval.BRAVO
BRAVO Bench
This study introduces BRAVO (Building Regulation Answering & Visual Observation) Bench, the first benchmark designed to evaluate the capability of Multimodal Large Language Models (MLLMs) to perform compliance checking based on BIM-derived scenes. BRAVO Bench integrates 26 BIM-scenes, 59 textual normative provisions, and 5 regulatory illustrations, producing 1505 question–answer pairs across four diagnostic layers: scene perception, scene understanding, rule… See the full description on the dataset page: https://huggingface.co/datasets/ritanibar/BRAVO.dotsocr-markdown-dataset
dotsocr_markdown_dataset
Dataset Description
This dataset contains training data for DotsOCR to convert document images directly to markdown format.
Training Objective
The model learns to:
Convert document images to clean markdown format
Preserve document structure and hierarchy
Extract all text content accurately
Use appropriate markdown formatting for different content types
Dataset Structure
Training samples: 798
Validation samples: 200
Total… See the full description on the dataset page: https://huggingface.co/datasets/rita1706/dotsocr-markdown-dataset.dotsocr_bank_statement_1K_v2
dotsocr_bank_statement_1K
Dataset Description
This dataset contains OCR and layout analysis training data formatted according to DotsOCR specifications by rednote-hilab.
DotsOCR Format Features
Proper Reading Order: Layout elements are sorted according to natural reading order (top to bottom, left to right)
Validated Categories: All categories conform to DotsOCR's specification: ['Caption', 'Footnote', 'Formula', 'List-item', 'Page-footer', 'Page-header'… See the full description on the dataset page: https://huggingface.co/datasets/rita1706/dotsocr_bank_statement_1K_v2.kitchenware_5000
Kitchen Utensils Dataset
Classes:
Bottle Opener
Bread Knife
Can Opener
Cup
Dessert Spoon
Fish Slice
Fork
Glass
Knife
Ladle
Masher
Peeler
Pizza Cutter
Plate
Serving Spoon
Spatula
Spoon
Tongs
Whisk
5184 pics
dotsocr_bank_statement_half
dotsocr_bank_statement_half
Dataset Description
This dataset contains OCR and layout analysis training data formatted according to DotsOCR specifications by rednote-hilab.
DotsOCR Format Features
Proper Reading Order: Layout elements are sorted according to natural reading order (top to bottom, left to right)
Validated Categories: All categories conform to DotsOCR's specification: ['Caption', 'Footnote', 'Formula', 'List-item', 'Page-footer', 'Page-header'… See the full description on the dataset page: https://huggingface.co/datasets/rita1706/dotsocr_bank_statement_half.imagenet_short_text_100_classes_x_100_samples
