datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
j2me-full-romset
🎮 J2ME Full Romset
A complete preservation archive of rare J2ME (Java Mobile) games — saved from digital extinction.
This collection was manually gathered, organized, and uploaded as part of the M5 Nostalgia Archive preservation project. Many of these titles are no longer available anywhere online and have been delisted by their original publishers.
📦 Contents
Folder
Files
Size
Roms/
456 .jar games
~303 MB
Photos/
1,381 screenshots
~153 MB… See the full description on the dataset page: https://huggingface.co/datasets/M5-Dev/j2me-full-romset.multilingual-coco
Multilingual Common Objects in Context (COCO) Dataset
This dataset is a collection of multiple language open-source captions of COCO dataset.
The split in this dataset is set according to Andrej Karpathy's split from dataset_coco.json file. The collection was created specifically for simplicity of use in training and evaluation pipeline by non-commercial and research purposes. The COCO images dataset is licensed under a Creative Commons Attribution 4.0 License.… See the full description on the dataset page: https://huggingface.co/datasets/romrawinjp/multilingual-coco.sas_sarageneral_light_curve_benchmark_dataset_collection_roman_simulated_variable_star_datasetROME
🏠Project Page & Leaderboard | 💻Code | 📄Paper | 🤗Data | 🤗Evaluation Response
This repository contains a visual reasoning benchmark named ROME from the paper FlagEval Findings Report: A Preliminary Evaluation of Large Reasoning Models on Automatically Verifiable Textual and Visual Questions.
ROME include 8 subtasks (281 high-quality questions in total). Each sample has been verified to ensure that images are necessary to answer correctly:
Academic
questions from college courses
Diagrams… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/ROME.romansr-data
specsr-roman training data — Roman grism spectra from OpenUniverse2024
Paired low-resolution Roman grism spectra and their noiseless ground-truth
SEDs, for training and evaluating spectral super-resolution. Built with
specsr-roman; the models trained on it are at
aryana-haghjoo/roman-spectral-superresolution.
Spectra
36,404
Unique galaxies
15,434
Redshift
0.036 – 3.005 (median 0.765; 35 % above z = 1)
Depth
AB(H158) 14.7 – 22.5
Pointings
5 visits × 18… See the full description on the dataset page: https://huggingface.co/datasets/aryana-haghjoo/romansr-data.multi30k
Multi30k Dataset
This dataset is a rearrangement version of the Multi30k dataset. The dataset was partially retrived from Multi30k original github.
How to use
This dataset can be downloaded from datasets library. train, validation, and test set are included in the dataset.
from datasets import load_dataset
dataset = load_dataset("romrawinjp/multi30k")
Reference
If you find this dataset beneficial, please directly cite to their incredible work.… See the full description on the dataset page: https://huggingface.co/datasets/romrawinjp/multi30k.mscoco
Common Objects in Context (COCO) Dataset
This dataset is English captions of COCO dataset.
The splits in this dataset is set according to Andrej Karpathy's split from dataset_coco.json file. The collection was created specifically for simplicity of use in training and evaluation pipeline by non-commercial and research purposes. The COCO images dataset is licensed under a Creative Commons Attribution 4.0 License.
Reference
@misc{lin2015microsoftcococommonobjects… See the full description on the dataset page: https://huggingface.co/datasets/romrawinjp/mscoco.orange_on_greenbackupromania_witdevanagari_and_roman_digits
Dataset Card for Dataset Name
The OCR Digits Dataset consists of 20,000 high-quality images of digit combinations captured under various conditions. This dataset is designed to support research in optical character recognition, particularly for multi-digit recognition tasks.
Dataset Details
Citation
BibTeX:
@dataset{SumitYadav2025OCRDigits,
author = {[Sumit Yadav]},
title = {OCR Digits Dataset: A Collection of 20,000 Multi-Digit(Roman and… See the full description on the dataset page: https://huggingface.co/datasets/rockerritesh/devanagari_and_roman_digits.COCO_keypointsMNIST-ResNet-Demo-DatatestImagesMr.Porter.Product.prices.Romania
Mr Porter web scraped data
About the website
Mr Porter operates in the expansive and rapidly evolving Ecommerce industry in EMEA, specifically within the fast-growing Romanian market. This sector has gained significant traction owing to the convergence of technology and commerce, encompassing a broad range of online business activities for products and services. Romania has emerged as a key player, given its robust digital infrastructure and tech-savvy population. The… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Mr.Porter.Product.prices.Romania.RuChartQA
RuChartQA
A Russian-language chart question answering benchmark for evaluating Vision-Language Models, with both synthetic and real-world evaluation sets.
Dataset summary
Split
Examples
Charts
Source
Synthetic ChartBasic
360
90 (×4 variants)
Generated
Synthetic ChartReasoning
480
120 (×4 variants)
Generated
Synthetic ChartPerception
360
90 (×4 variants)
Generated
ChartReal
242 QA
96 charts
Rosstat, Bank of Russia (PDF)
Total
1442 QA
396 unique charts… See the full description on the dataset page: https://huggingface.co/datasets/romath/RuChartQA.romanian-driving-examro_mmmu
Dataset Description
MMMU is a a new benchmark designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning.
Here we provide the Romanian translation of MMUU, translated with gpt-4.1-mini. This dataset is used as a benchmark and is part of the evaluation protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_mmmu.roman_numeralsGucci.Product.prices.Romania
Gucci web scraped data
About the website
Operating within the luxury fashion industry, Gucci marks its prominent presence in the EMEA region, specifically in Romania. The luxury fashion industry in Romania has shown considerable growth, owing to increased consumer spending and changing lifestyle trends. The industry is characterized by high-income consumers with a taste for luxury fashion products displaying their status and personality. The industry, driven by premium… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Gucci.Product.prices.Romania.pokemon-captionsRoMMathromeo-rosete
Dataset Card for Dataset Name
https://huggingface.co/docs/datasets/index
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
dataset = dataset.cast_column("audio", Audio(sampling_rate=16000))
dataset[0]["audio"]
Dataset Description
import { pipeline } from '@huggingface/transformers';
// Allocate a pipeline for sentiment-analysis
const pipe = await pipeline('sentiment-analysis');… See the full description on the dataset page: https://huggingface.co/datasets/bombastictranz/romeo-rosete.RoMemes
Dataset Description
RoMemes is a dataset of Romanian language memes, collected from public social media platforms.
This dataset is used as a benchmark and is part of the evaluation protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation
@article{puaics2024romemes,
title={RoMemes: A multimodal meme corpus for the Romanian language},
author={P{\u{a}}i{\c{s}}, Vasile and… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/RoMemes.ghai_romaine_detection
Ghai Romaine Detection
A dataset for object detection of romaine. The dataset contains 500 images with 2,323 bounding box annotations across 1 category.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
https://github.com/AxisAg/GHAIDatasets/blob/main/datasets/romaine.md
ro_mmstar
Dataset Description
MMStar an elite vision-indispensable multi-modal benchmark comprising 1,500 challenge samples meticulously selected by humans.
Here we provide the Romanian translation of MMStar, translated with gpt-4.1-mini. This dataset is used as a benchmark and is part of the evaluation protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation
@article{chen2024we… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_mmstar.SquareNut_D0
romantic_topro_mme
Dataset Description
MME is a benchmark that measures both perception and cognition abilities on a total of 14 subtasks.
Here we provide the Romanian translation of SEED-Bench-2, translated with gpt-4.1-mini. This dataset is used as a benchmark and is part of the evaluation protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation
@article{fu2026mme,
title={Mme: A comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_mme.dataset_chromosome
