CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01roman-bushuiev /MassSpecGym MassSpecGym provides a dataset and benchmark for the discovery and identification of new molecules from tandem mass spectrometry (MS/MS) spectra. The provided challenges abstract the process of scientific discovery of new molecules from biological and environmental samples into well-defined machine learning problems. Papers MassSpecGym in the Wild: Uncovering and Correcting Evaluation Pitfalls in AI-Driven Molecule Discovery (2025): Paper Link MassSpecGym: A benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/roman-bushuiev/MassSpecGym.tabularother100K<n<1M23 likes3.5k downloads2mo agoHugging Face02Mavkif /Roman-Urdu-Parl-split Roman Urdu Parallel Dataset - Split This dataset is just another version of Roman-Urdu-Parl dataset split into train, validation and test set properly. Details follow below. This repository contains a split version of the Roman-Urdu Parallel Dataset (Roman-Urdu-Parl) structured specifically to facilitate fair evaluation in machine transliteration tasks between Urdu and Roman-Urdu. The Roman-Urdu language lacks standard orthography, leading to a wide range of transliteration… See the full description on the dataset page: https://huggingface.co/datasets/Mavkif/Roman-Urdu-Parl-split.texttranslation1M<n<10M5 likes214 downloads2y agoHugging Face03dejp3 /roman-coins-PASThis dataset is a csv file of python scraped data from the JSON API of the Portable Antiquities Scheme website tabular100K<n<1M0 likes124 downloads11mo agoHugging Face04Romit2004 /LinuxCommandstext10K<n<100K28 likes117 downloads3y agoHugging Face05Roman-Kpro /russian-business-registries Russian Business Registries — Aggregated Statistics Aggregated, ready-to-analyse slices of Russian state registers. Every figure comes from an official open-data source; nothing here is modelled, imputed or estimated. Individual companies are not published — only aggregates, with one deliberate exception described below. Собрано из открытых данных российских госреестров. Все цифры — из официальных источников, без моделирования и досчётов. Публикуются агрегаты, не сведения об… See the full description on the dataset page: https://huggingface.co/datasets/Roman-Kpro/russian-business-registries.tabulartabular-classification10K<n<100K0 likes109 downloads9d agoHugging Face06Romjiik /Russian_bank_reviews Dataset Card for bank reviews dataset Dataset Summary The dataset is collected from the banki.ru website. It contains customer reviews of various banks. In total, the dataset contains 12399 reviews. The dataset is suitable for sentiment classification. The dataset contains this fields - bank name, username, review title, review text, review time, number of views, number of comments, review rating set by the user, as well as ratings for special categories… See the full description on the dataset page: https://huggingface.co/datasets/Romjiik/Russian_bank_reviews.tabulartext-classification10K<n<100K7 likes106 downloads3y agoHugging Face07sk-community /romanized_hindi Romanized Hindi Dataset Dataset Description The Romanized Hindi Dataset is a collection of Hindi text paired with its Romanized (Latin script) representation. It has been created by combining multiple sources, including open datasets, synthetic generation, and rule-based transliteration methods. The dataset is designed for training and evaluating Hindi↔Roman transliteration models. Language(s): Hindi, Romanized Hindi Size: ~1.82M rows License: MIT (check with source… See the full description on the dataset page: https://huggingface.co/datasets/sk-community/romanized_hindi.text1M<n<10M0 likes106 downloads1y agoHugging Face08radool /romanian-name-days Romanian Name Days and Holidays Zile onomastice și sărbători românești — the Romanian name-day calendar as structured data. In Romania, ziua onomastică — the feast day of the saint whose name you bear — is widely celebrated, often more than a birthday. Until now this information existed online only as HTML pages built for human readers. This is the machine-readable version. Published by trends.ro. Dataset summary Names 86 (46 masculine, 40 feminine)… See the full description on the dataset page: https://huggingface.co/datasets/radool/romanian-name-days.textquestion-answeringn<1K1 likes94 downloads1mo agoHugging Face09Romaindcrl /ohlcv_ethusdttabular100K<n<1M2 likes77 downloads2y agoHugging Face10Roman190928 /LargeAlgebra Use how ever the heck u want, idc about it, im nice and it took me like 1 hr to make this, so not very long. 60 million lines (1million per file) For matting is like this #equation | answer | reasoning problem | answer | why that answer works id recommend doing a train test split of 75/25 to avoid overfitting Unliscensed Let me know if you'd like another db. text10M<n<100M0 likes70 downloads11mo agoHugging Face11WebSEM-ai /agent-discoverability-ado-score-romania Agent Discoverability (ADO Score) — Romania, September 2026 130 Romanian domains probed for A2A Agent Cards, MCP discovery, llms.txt, schema.org and Wikidata. Zero Agent Cards; mean ADO Score 17/100. Raw data, scripts and scoring spec, CC BY 4.0. Canonical study (analysis, charts, interpretation): Romanian · English What this is On 8 September 2026 a standard-library Python probe (published) requested, for each of 130 domains, the homepage without JavaScript… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/agent-discoverability-ado-score-romania.tabularn<1K0 likes69 downloads14d agoHugging Face12Telugu-LLM-Labs /telugu_alpaca_yahma_cleaned_filtered_romanizedtext10K<n<100K19 likes59 downloads3y agoHugging Face13fatymahaly /Roman-Urdu-Sentiment-Dataset Roman Urdu Sentiment Dataset This repository contains a curated dataset of Roman Urdu text collected from social media interactions, comments, and daily online conversations. Each text entry is paired with a sentiment label for Natural Language Processing (NLP) tasks such as sentiment analysis. Dataset Structure The dataset is formatted in comma-separated values (.csv) with the following columns: text: The sentence, phrase, or social media comment written in… See the full description on the dataset page: https://huggingface.co/datasets/fatymahaly/Roman-Urdu-Sentiment-Dataset.textn<1K0 likes57 downloads14d agoHugging Face14craciuncg /RoMedQA_v1 Description This is RoMedQA, a dataset that amounts to 4,127 single-choice questions regarding the medical field in the Romanian language. The dataset consists of advanced biology questions used in entrance examinations in medical schools in Romania. Each question has five possible answer choices, numbered from 1 to 5, with only one correct answer. Loading For loading the dataset, you can simply proceed as follows: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/craciuncg/RoMedQA_v1.textquestion-answering1K<n<10K1 likes47 downloads2y agoHugging Face15craciuncg /RoMedQA_v2textquestion-answering1K<n<10K1 likes44 downloads2y agoHugging Face16Khubaib01 /roman-urdu-sentiment-embeddings Roman Urdu Sentiment Embeddings Dataset Overview This repository contains a research-grade Roman Urdu Sentiment Analysis dataset released as anonymized sentence embeddings. Roman Urdu is a low-resource language with high linguistic variability, heavy slang usage, and frequent code-mixing with English. This dataset is curated to support robust sentiment analysis research while ensuring strict privacy preservation. Raw text data has not been released. Instead, all messages… See the full description on the dataset page: https://huggingface.co/datasets/Khubaib01/roman-urdu-sentiment-embeddings.text10K<n<100K2 likes40 downloads9mo agoHugging Face17Khubaib01 /RomanUrdu-NLP-Sentiment-Corpus RomanUrdu-NLP-Sentiment-Corpus Largest Open-Source Roman Urdu Sentiment Dataset with Slang Robustness Overview This repository presents the largest publicly available Roman Urdu sentiment analysis dataset, containing 134,052 labeled text samples collected from chats and social media platforms. The dataset is designed to be: Robust to slang and informal Roman Urdu High-quality through LLM-assisted labeling and human validation Balanced across sentiment classes… See the full description on the dataset page: https://huggingface.co/datasets/Khubaib01/RomanUrdu-NLP-Sentiment-Corpus.tabulartext-classification100K<n<1M2 likes40 downloads7mo agoHugging Face18Qasim522 /Roman-Urdu-Parl-split Roman Urdu Parallel Dataset - Split This dataset is just another version of Roman-Urdu-Parl dataset split into train, validation and test set properly. Details follow below. This repository contains a split version of the Roman-Urdu Parallel Dataset (Roman-Urdu-Parl) structured specifically to facilitate fair evaluation in machine transliteration tasks between Urdu and Roman-Urdu. The Roman-Urdu language lacks standard orthography, leading to a wide range of transliteration… See the full description on the dataset page: https://huggingface.co/datasets/Qasim522/Roman-Urdu-Parl-split.texttranslation1M<n<10M0 likes38 downloads9mo agoHugging Face19WebSEM-ai /ai-visibility-romania-luxury-jewelry AI Visibility — Romania's Luxury Jewelry (July 2026) We asked ten large language models the same question, word for word. We got 29 different brands across 50 available positions, and no brand appeared in all ten lists. Canonical study (analysis, charts, interpretation): Romanian · English The prompt Which luxury jewelry brands from Romania do you recommend for wedding bands and engagement rings? Give me a top 5, with a short argument for each and the sources you… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/ai-visibility-romania-luxury-jewelry.tabularn<1K1 likes29 downloads2mo agoHugging Face20Telugu-LLM-Labs /telugu_teknium_GPTeacher_general_instruct_filtered_romanizedtext10K<n<100K15 likes28 downloads3y agoHugging Face21Madu786 /multilingual-urdu-romanurdu-arabic-english-sentiment Multilingual Sentiment Classification Dataset (EN, UR, Roman UR, AR) Overview This dataset is a clean and balanced multilingual sentiment classification dataset covering four languages: English Urdu Roman Urdu Arabic The dataset is designed to support sentiment analysis and text classification tasks, especially for low-resource languages such as Urdu and Roman Urdu. Sentiment Classes Each text sample belongs to one of the following sentiment categories:… See the full description on the dataset page: https://huggingface.co/datasets/Madu786/multilingual-urdu-romanurdu-arabic-english-sentiment.text100K<n<1M0 likes28 downloads9mo agoHugging Face22Speech-data /Romanian-Speech-Dataset 🎧 Romanian Speech Dataset The Romanian Speech Dataset is a high-quality speech audio dataset designed to support AI and machine learning workflows with diverse and well-structured audio data. It includes 117 hours of recorded speech data across 878 files, delivered in MP3 and WAV formats, with a total size of 188 MB. This carefully curated audio dataset provides balanced and representative voice data, with 54% male and 46% female speakers, and age distribution spanning 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Romanian-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes27 downloads6mo agoHugging Face23DenSeduct /Russian_Romantic_Dialogue_Dataset 💬 Russian Romantic Dialogue Dataset — Real AI-to-Human Conversations 🧩 Dataset Summary This dataset contains annotated Russian-language romantic dialogues produced by a proprietary AI dating assistant operating in production across multiple platforms. Each dialogue is a real interaction between an AI system and a real woman responding in natural conditions. The women respond naturally, producing authentic emotional dynamics, trust signals, resistance patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/DenSeduct/Russian_Romantic_Dialogue_Dataset.tabularn<1K1 likes27 downloads4mo agoHugging Face24deshanksuman /Swabhasha_RomanizedSinhala_Dataset Model Card for Model ID This Repo is about Romanized Sinhala to Sinhala Transliteration using the Ngram and Rule Base Model. Model Description This dataset is capable of handling short-hand typing(Adhoc Transliteration). eg Input: khmda Output : කොහොමද If you are using this work: Kindly cite : T. G. D. K. Sumanathilaka, R. Weerasinghe and Y. H. P. P. Priyadarshana, "Swa-Bhasha: Romanized Sinhala to Sinhala Reverse Transliteration using a Hybrid Approach," 2023 3rd… See the full description on the dataset page: https://huggingface.co/datasets/deshanksuman/Swabhasha_RomanizedSinhala_Dataset.text1M<n<10M2 likes25 downloads1y agoHugging Face25Romoamigo /oop-bad-code-to-good-code-cpptext10K<n<100K0 likes25 downloads1y agoHugging Face26dipeshch71 /nepaliflow-romanized-nepali-to-devanagari-dataset NepaliFlow Romanized Nepali to Devanagari Dataset This dataset contains instruction-style examples for converting Romanized Nepali words into Nepali Devanagari script. Task The task is to convert a Romanized Nepali word into its Devanagari form while returning only the Devanagari output. Columns prompt: instruction asking the model to convert a Romanized Nepali word into Devanagari completion: expected Nepali Devanagari output Size… See the full description on the dataset page: https://huggingface.co/datasets/dipeshch71/nepaliflow-romanized-nepali-to-devanagari-dataset.text10K<n<100K1 likes25 downloads3mo agoHugging Face27WebSEM-ai /ai-search-visibility-romania-electronics-market AI Search Visibility — Romania's Electronics & IT Market (August 2026) 18 brand-free purchase questions × 5 AI engines = 87 answers. 86 of them name a major retailer. Position, not presence, decides the market. Raw data CC BY 4.0. Canonical study (analysis, charts, interpretation): Romanian · English What this is Eighteen real purchase questions were put to ChatGPT, Google Gemini, Perplexity, Google AI Mode and Google AI Overviews, in Romanian, from Romania, in… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/ai-search-visibility-romania-electronics-market.tabularn<1K0 likes25 downloads2mo agoHugging Face28DGurgurov /romanian_sa Sentiment Analysis Data for the Romanian Language Dataset Description: This dataset contains a sentiment analysis dataset from Tache et al. (2021). Data Structure: The data was used for the project on improving word embeddings with graph knowledge for Low Resource Languages. Citation: @inproceedings{tache-etal-2021-clustering, title = "Clustering Word Embeddings with Self-Organizing Maps. Application on {L}a{R}o{S}e{D}a - A Large {R}omanian Sentiment Data Set", author =… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/romanian_sa.texttext-classification10K<n<100K0 likes22 downloads2y agoHugging Face29sanujen /Legacy-Font-and-Romanized-Tamil-Corpustexttext-classification10K<n<100K0 likes22 downloads1y agoHugging Face30Aimlab /Sentiment-Analysis-Roman-Urdutext10K<n<100K6 likes21 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.