datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ita-eval-resultsThis dataset page contains numerical results to populate the ItaEval Leaderboard.
It is intended for internal use only.
Adding new Results
Once new results are computed, to include them in the leadboard it only takes to save them here in a results.json file. Note, though, that the file must contain several fields as specified in the leaderboard app. Use another result.json file for reference.
Adding info of a new Model
Since we do not support automatic submission and… See the full description on the dataset page: https://huggingface.co/datasets/RiTA-nlp/ita-eval-results.masscount-cf
MassCount-CF
Synthetic counterfactual counting corpus for identifying an additive
neural-mass cardinality coordinate in vision-language models (NMCA).
Companion dataset to "Counting Requires Mass: An Algebraic and Causal Account
of Numerosity in Vision-Language Models."
Corpus version: masscount-cf-1.0.0
Master scenes: 99,989
Delivered images (this upload): 326,623
Object instances: 28,617,018
Structure
Scene graphs are the durable artifact; pixels are regenerable… See the full description on the dataset page: https://huggingface.co/datasets/Ritabrata04/masscount-cf.icub_sim_dataset_t2_smooth_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "icub",
"total_episodes": 90,
"total_frames": 9935,
"total_tasks": 5,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ritamota/icub_sim_dataset_t2_smooth_lerobot.clinical-synthetic-text-kg
Data Description
We release the synthetic data generated using the method described in the paper Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models
(ACL 2024 Findings). The external knowledge we use is based on external knowledge graphs.
Generated Datasets
The original train/validation/test data, and the generated synthetic training data are listed as follows. For each dataset, we generate 5000 synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ritaranx/clinical-synthetic-text-kg.icub_sim_dataset_t4_smooth_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "icub",
"total_episodes": 90,
"total_frames": 9216,
"total_tasks": 90,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ritamota/icub_sim_dataset_t4_smooth_lerobot.clinical-synthetic-text-llm
Data Description
We release the synthetic data generated using the method described in the paper Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models
(ACL 2024 Findings). The external knowledge we use is based on LLM-generated topics and writing styles.
Generated Datasets
The original train/validation/test data, and the generated synthetic training data are listed as follows. For each dataset, we generate 5000… See the full description on the dataset page: https://huggingface.co/datasets/ritaranx/clinical-synthetic-text-llm.icub_sim_dataset_t3_smooth_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "icub",
"total_episodes": 90,
"total_frames": 4000,
"total_tasks": 90,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ritamota/icub_sim_dataset_t3_smooth_lerobot.icub_sim_dataset_t1_smooth_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "icub",
"total_episodes": 90,
"total_frames": 12632,
"total_tasks": 64,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ritamota/icub_sim_dataset_t1_smooth_lerobot.biomni-datalakerita_ainsworth_sakurasounopetnakanojo
Dataset of Rita Ainsworth (Sakurasou no Pet na Kanojo)
This is the dataset of Rita Ainsworth (Sakurasou no Pet na Kanojo), containing 105 images and their tags.
The core tags of this character are blonde_hair, long_hair, blue_eyes, bangs, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
List of Packages
Name
Images
Size… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/rita_ainsworth_sakurasounopetnakanojo.icub_sim_dataset_t5_smooth_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "icub",
"total_episodes": 90,
"total_frames": 10716,
"total_tasks": 90,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ritamota/icub_sim_dataset_t5_smooth_lerobot.rit-arc-challenge
ARC-Challenge Multilingual
This repository contains the multilingual ARC-Challenge benchmark release from
Recovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets
(project page, arXiv:2602.22207).
It reorganizes the per-language datasets from the
INSAIT-Institute/multilingual-benchmarks
collection into one Hugging Face dataset repository, with one config/subset per language.
ARC-Challenge contains grade-school science questions that… See the full description on the dataset page: https://huggingface.co/datasets/INSAIT-Institute/rit-arc-challenge.ai2_arc_ita
Dataset Card for Ai2 ARC (ita)
This dataset is a machine-translated version of Ai2 ARC into Italian.
Licensed under CC-BY 4.0
Translated with TowerInstruct-7B-v0.2
More details and code used for translation will follow shortly.
The rest of the page is WIP :)
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/RiTA-nlp/ai2_arc_ita.Aegis-AI-Content-Safety-Dataset-2.0
🛡️ Nemotron Content Safety Dataset V2
The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1.
To curate the dataset, we use the HuggingFace version of human preference data about… See the full description on the dataset page: https://huggingface.co/datasets/ritatai727/Aegis-AI-Content-Safety-Dataset-2.0.counterfactual_dataset_20_classes_x_100_samplesrita_henschel_alicegearaegisexpansion
Dataset of Rita Henschel
This is the dataset of Rita Henschel, containing 50 images and their tags.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
Name
Images
Download
Description
raw
50
Download
Raw data with meta information.
raw-stage3
121
Download
3-stage cropped raw data with meta information.
raw-stage3-eyes
139
Download
3-stage cropped (with eye-focus) raw… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/rita_henschel_alicegearaegisexpansion.UINAUILGeNTE_ita-evalThis dataset adapts the GeNTE dataset to be run on ItaEval.
In particular, we selected the five initial instances from the dataset and used them as few-shot examples ("train" split in this release).
Please check the original dataset for more details on GeNTE. If you use the dataset, please consider citing the paper:
@inproceedings{piergentili-etal-2023-hi,
title = "Hi Guys or Hi Folks? Benchmarking Gender-Neutral Machine Translation with the {G}e{NTE} Corpus",
author = "Piergentili… See the full description on the dataset page: https://huggingface.co/datasets/RiTA-nlp/GeNTE_ita-eval.sold-alpacaITALICITALIC is a dataset of Italian audio recordings and contains annotation for utterance transcripts and associated intents.
The ITALIC dataset was created through a custom web platform, utilizing both native and non-native Italian speakers as participants.
The participants were required to record themselves while reading a randomly sampled short text from the MASSIVE dataset.icub_sim_dataset_t3_simple_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "icub",
"total_episodes": 90,
"total_frames": 19830,
"total_tasks": 5,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ritamota/icub_sim_dataset_t3_simple_lerobot.Ritammrag-eval
mmrag-eval
Benchmark dataset for evaluating grounding quality in multimodal Retrieval-Augmented Generation (RAG) systems.
Standard benchmarks measure whether a RAG system retrieves the right document. mmrag-eval measures whether the system's generated answer faithfully reflects what is actually shown in the retrieved image — catching hallucination and retrieval redundancy that retrieval metrics alone cannot detect.
Dataset Summary
198 annotated image–query pairs… See the full description on the dataset page: https://huggingface.co/datasets/ritaban-b/mmrag-eval.truthful_qa_ita
Dataset Card TruthfulQA (ita)
This dataset is a machine-translated version of truthful_qa into Italian.
Licensed under CC-BY 4.0
Translated with TowerInstruct-7B-v0.2
More details and code used for translation will follow shortly.
The rest of the page is WIP :)
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/RiTA-nlp/truthful_qa_ita.BRAVO
BRAVO Bench
This study introduces BRAVO (Building Regulation Answering & Visual Observation) Bench, the first benchmark designed to evaluate the capability of Multimodal Large Language Models (MLLMs) to perform compliance checking based on BIM-derived scenes. BRAVO Bench integrates 26 BIM-scenes, 59 textual normative provisions, and 5 regulatory illustrations, producing 1505 question–answer pairs across four diagnostic layers: scene perception, scene understanding, rule… See the full description on the dataset page: https://huggingface.co/datasets/ritanibar/BRAVO.dotsocr-markdown-dataset
dotsocr_markdown_dataset
Dataset Description
This dataset contains training data for DotsOCR to convert document images directly to markdown format.
Training Objective
The model learns to:
Convert document images to clean markdown format
Preserve document structure and hierarchy
Extract all text content accurately
Use appropriate markdown formatting for different content types
Dataset Structure
Training samples: 798
Validation samples: 200
Total… See the full description on the dataset page: https://huggingface.co/datasets/rita1706/dotsocr-markdown-dataset.20260504RL_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 32,
"total_frames": 12528,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/RITAHuang/20260504RL_test.hellaswag_itaRita_Black_white_mix
📚 Rita: Class 8 Math Video Dataset (Raw)
Primary Training Data for Zulense Z1
This dataset contains raw video footage of Class 8 Mathematics lectures (Indian Curriculum/NCERT), categorized by teaching medium (Blackboard vs. Whiteboard) and teacher presence.
This data is used to train the Zulense Imagination Engine to understand the difference between:
Teacher Dynamics: How a teacher moves and gestures.
Content Logic: How equations and diagrams appear on a board over time.… See the full description on the dataset page: https://huggingface.co/datasets/ProgramerSalar/Rita_Black_white_mix.sinhalese-sentiment-alpaca
