datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MassSpecGym
MassSpecGym provides a dataset and benchmark for the discovery and identification of new molecules from tandem mass spectrometry (MS/MS) spectra. The provided challenges abstract the process of scientific discovery of new molecules from biological and environmental samples into well-defined machine learning problems.
Papers
MassSpecGym in the Wild: Uncovering and Correcting Evaluation Pitfalls in AI-Driven Molecule Discovery (2025): Paper Link
MassSpecGym: A benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/roman-bushuiev/MassSpecGym.Roman-Urdu-Parl-split
Roman Urdu Parallel Dataset - Split
This dataset is just another version of Roman-Urdu-Parl dataset split into train, validation and test set properly. Details follow below.
This repository contains a split version of the Roman-Urdu Parallel Dataset (Roman-Urdu-Parl) structured specifically to facilitate fair evaluation in machine transliteration tasks between Urdu and Roman-Urdu.
The Roman-Urdu language lacks standard orthography, leading to a wide range of transliteration… See the full description on the dataset page: https://huggingface.co/datasets/Mavkif/Roman-Urdu-Parl-split.roman-coins-PASThis dataset is a csv file of python scraped data from the JSON API of the Portable Antiquities Scheme website
LinuxCommandsrussian-business-registries
Russian Business Registries — Aggregated Statistics
Aggregated, ready-to-analyse slices of Russian state registers. Every figure
comes from an official open-data source; nothing here is modelled, imputed or
estimated. Individual companies are not published — only aggregates, with one
deliberate exception described below.
Собрано из открытых данных российских госреестров. Все цифры — из официальных
источников, без моделирования и досчётов. Публикуются агрегаты, не сведения
об… See the full description on the dataset page: https://huggingface.co/datasets/Roman-Kpro/russian-business-registries.Russian_bank_reviews
Dataset Card for bank reviews dataset
Dataset Summary
The dataset is collected from the banki.ru website.
It contains customer reviews of various banks. In total, the dataset contains 12399 reviews.
The dataset is suitable for sentiment classification.
The dataset contains this fields - bank name, username, review title, review text, review time, number of views,
number of comments, review rating set by the user, as well as ratings for special categories… See the full description on the dataset page: https://huggingface.co/datasets/Romjiik/Russian_bank_reviews.romanized_hindi
Romanized Hindi Dataset
Dataset Description
The Romanized Hindi Dataset is a collection of Hindi text paired with its Romanized (Latin script) representation.
It has been created by combining multiple sources, including open datasets, synthetic generation, and rule-based transliteration methods.
The dataset is designed for training and evaluating Hindi↔Roman transliteration models.
Language(s): Hindi, Romanized Hindi
Size: ~1.82M rows
License: MIT (check with source… See the full description on the dataset page: https://huggingface.co/datasets/sk-community/romanized_hindi.romanian-name-days
Romanian Name Days and Holidays
Zile onomastice și sărbători românești — the Romanian name-day calendar as
structured data.
In Romania, ziua onomastică — the feast day of the saint whose name you bear —
is widely celebrated, often more than a birthday. Until now this information
existed online only as HTML pages built for human readers. This is the
machine-readable version.
Published by trends.ro.
Dataset summary
Names
86 (46 masculine, 40 feminine)… See the full description on the dataset page: https://huggingface.co/datasets/radool/romanian-name-days.ohlcv_ethusdtLargeAlgebra
Use how ever the heck u want, idc about it, im nice and it took me like 1 hr to make this, so not very long.
60 million lines (1million per file)
For matting is like this
#equation | answer | reasoning
problem | answer | why that answer works
id recommend doing a train test split of 75/25 to avoid overfitting
Unliscensed
Let me know if you'd like another db.
agent-discoverability-ado-score-romania
Agent Discoverability (ADO Score) — Romania, September 2026
130 Romanian domains probed for A2A Agent Cards, MCP discovery, llms.txt, schema.org and Wikidata. Zero Agent Cards; mean ADO Score 17/100. Raw data, scripts and scoring spec, CC BY 4.0.
Canonical study (analysis, charts, interpretation):
Romanian ·
English
What this is
On 8 September 2026 a standard-library Python probe (published) requested, for each of 130 domains, the homepage without JavaScript… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/agent-discoverability-ado-score-romania.telugu_alpaca_yahma_cleaned_filtered_romanizedRoman-Urdu-Sentiment-Dataset
Roman Urdu Sentiment Dataset
This repository contains a curated dataset of Roman Urdu text collected from social media interactions, comments, and daily online conversations. Each text entry is paired with a sentiment label for Natural Language Processing (NLP) tasks such as sentiment analysis.
Dataset Structure
The dataset is formatted in comma-separated values (.csv) with the following columns:
text: The sentence, phrase, or social media comment written in… See the full description on the dataset page: https://huggingface.co/datasets/fatymahaly/Roman-Urdu-Sentiment-Dataset.RoMedQA_v1
Description
This is RoMedQA, a dataset that amounts to 4,127 single-choice questions regarding the medical field in the Romanian language.
The dataset consists of advanced biology questions used in entrance examinations in medical schools in Romania.
Each question has five possible answer choices, numbered from 1 to 5, with only one correct answer.
Loading
For loading the dataset, you can simply proceed as follows:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/craciuncg/RoMedQA_v1.RoMedQA_v2roman-urdu-sentiment-embeddings
Roman Urdu Sentiment Embeddings Dataset
Overview
This repository contains a research-grade Roman Urdu Sentiment Analysis dataset released as anonymized sentence embeddings. Roman Urdu is a low-resource language with high linguistic variability, heavy slang usage, and frequent code-mixing with English. This dataset is curated to support robust sentiment analysis research while ensuring strict privacy preservation.
Raw text data has not been released. Instead, all messages… See the full description on the dataset page: https://huggingface.co/datasets/Khubaib01/roman-urdu-sentiment-embeddings.RomanUrdu-NLP-Sentiment-Corpus
RomanUrdu-NLP-Sentiment-Corpus
Largest Open-Source Roman Urdu Sentiment Dataset with Slang Robustness
Overview
This repository presents the largest publicly available Roman Urdu sentiment analysis dataset, containing 134,052 labeled text samples collected from chats and social media platforms. The dataset is designed to be:
Robust to slang and informal Roman Urdu
High-quality through LLM-assisted labeling and human validation
Balanced across sentiment classes… See the full description on the dataset page: https://huggingface.co/datasets/Khubaib01/RomanUrdu-NLP-Sentiment-Corpus.Roman-Urdu-Parl-split
Roman Urdu Parallel Dataset - Split
This dataset is just another version of Roman-Urdu-Parl dataset split into train, validation and test set properly. Details follow below.
This repository contains a split version of the Roman-Urdu Parallel Dataset (Roman-Urdu-Parl) structured specifically to facilitate fair evaluation in machine transliteration tasks between Urdu and Roman-Urdu.
The Roman-Urdu language lacks standard orthography, leading to a wide range of transliteration… See the full description on the dataset page: https://huggingface.co/datasets/Qasim522/Roman-Urdu-Parl-split.ai-visibility-romania-luxury-jewelry
AI Visibility — Romania's Luxury Jewelry (July 2026)
We asked ten large language models the same question, word for word. We got
29 different brands across 50 available positions, and no brand appeared in
all ten lists.
Canonical study (analysis, charts, interpretation):
Romanian ·
English
The prompt
Which luxury jewelry brands from Romania do you recommend for wedding bands
and engagement rings? Give me a top 5, with a short argument for each and the
sources you… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/ai-visibility-romania-luxury-jewelry.telugu_teknium_GPTeacher_general_instruct_filtered_romanizedmultilingual-urdu-romanurdu-arabic-english-sentiment
Multilingual Sentiment Classification Dataset (EN, UR, Roman UR, AR)
Overview
This dataset is a clean and balanced multilingual sentiment classification dataset
covering four languages:
English
Urdu
Roman Urdu
Arabic
The dataset is designed to support sentiment analysis and text classification
tasks, especially for low-resource languages such as Urdu and Roman Urdu.
Sentiment Classes
Each text sample belongs to one of the following sentiment categories:… See the full description on the dataset page: https://huggingface.co/datasets/Madu786/multilingual-urdu-romanurdu-arabic-english-sentiment.Romanian-Speech-Dataset
🎧 Romanian Speech Dataset
The Romanian Speech Dataset is a high-quality speech audio dataset designed to support AI and machine learning workflows with diverse and well-structured audio data. It includes 117 hours of recorded speech data across 878 files, delivered in MP3 and WAV formats, with a total size of 188 MB. This carefully curated audio dataset provides balanced and representative voice data, with 54% male and 46% female speakers, and age distribution spanning 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Romanian-Speech-Dataset.Russian_Romantic_Dialogue_Dataset
💬 Russian Romantic Dialogue Dataset — Real AI-to-Human Conversations
🧩 Dataset Summary
This dataset contains annotated Russian-language romantic dialogues produced by a proprietary AI dating assistant operating in production across multiple platforms. Each dialogue is a real interaction between an AI system and a real woman responding in natural conditions.
The women respond naturally, producing authentic emotional dynamics, trust signals, resistance patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/DenSeduct/Russian_Romantic_Dialogue_Dataset.Swabhasha_RomanizedSinhala_Dataset
Model Card for Model ID
This Repo is about Romanized Sinhala to Sinhala Transliteration using the Ngram and Rule Base Model.
Model Description
This dataset is capable of handling short-hand typing(Adhoc Transliteration).
eg
Input: khmda
Output : කොහොමද
If you are using this work:
Kindly cite :
T. G. D. K. Sumanathilaka, R. Weerasinghe and Y. H. P. P. Priyadarshana, "Swa-Bhasha: Romanized Sinhala to Sinhala Reverse Transliteration using a Hybrid Approach," 2023 3rd… See the full description on the dataset page: https://huggingface.co/datasets/deshanksuman/Swabhasha_RomanizedSinhala_Dataset.oop-bad-code-to-good-code-cppnepaliflow-romanized-nepali-to-devanagari-dataset
NepaliFlow Romanized Nepali to Devanagari Dataset
This dataset contains instruction-style examples for converting Romanized Nepali words into Nepali Devanagari script.
Task
The task is to convert a Romanized Nepali word into its Devanagari form while returning only the Devanagari output.
Columns
prompt: instruction asking the model to convert a Romanized Nepali word into Devanagari
completion: expected Nepali Devanagari output
Size… See the full description on the dataset page: https://huggingface.co/datasets/dipeshch71/nepaliflow-romanized-nepali-to-devanagari-dataset.ai-search-visibility-romania-electronics-market
AI Search Visibility — Romania's Electronics & IT Market (August 2026)
18 brand-free purchase questions × 5 AI engines = 87 answers. 86 of them name a major retailer. Position, not presence, decides the market. Raw data CC BY 4.0.
Canonical study (analysis, charts, interpretation):
Romanian ·
English
What this is
Eighteen real purchase questions were put to ChatGPT, Google Gemini, Perplexity, Google AI Mode and Google AI Overviews, in Romanian, from Romania, in… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/ai-search-visibility-romania-electronics-market.romanian_sa
Sentiment Analysis Data for the Romanian Language
Dataset Description:
This dataset contains a sentiment analysis dataset from Tache et al. (2021).
Data Structure:
The data was used for the project on improving word embeddings with graph knowledge for Low Resource Languages.
Citation:
@inproceedings{tache-etal-2021-clustering,
title = "Clustering Word Embeddings with Self-Organizing Maps. Application on {L}a{R}o{S}e{D}a - A Large {R}omanian Sentiment Data Set",
author =… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/romanian_sa.Legacy-Font-and-Romanized-Tamil-CorpusSentiment-Analysis-Roman-Urdu
