CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RomanJordansky /BOT_JORDANS-storagegatedtextn<1K0 likes4.6k downloads8d agoHugging Face02roman-bushuiev /MassSpecGym MassSpecGym provides a dataset and benchmark for the discovery and identification of new molecules from tandem mass spectrometry (MS/MS) spectra. The provided challenges abstract the process of scientific discovery of new molecules from biological and environmental samples into well-defined machine learning problems. Papers MassSpecGym in the Wild: Uncovering and Correcting Evaluation Pitfalls in AI-Driven Molecule Discovery (2025): Paper Link MassSpecGym: A benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/roman-bushuiev/MassSpecGym.tabularother100K<n<1M23 likes3.5k downloads2mo agoHugging Face03Romoamigo /SWE-Bench-MultilingualC_CPPFileteredtextn<1K0 likes1.8k downloads1y agoHugging Face04Romoamigo /SWE-Bench-MultilingualC_CPPFiletered_newtextn<1K0 likes1.6k downloads1y agoHugging Face0515juneee /romanian-corpus Romanian Text Corpus A comprehensive, high-quality Romanian text corpus for language model pretraining. Built by collecting and cleaning text from five Romanian-language sources. Dataset Summary Total documents: 19,886,412 Estimated tokens: ~20.8B Language: Romanian (ro) Format: Parquet (zstd compressed) Source Breakdown Source Documents mC4 16,875,310 OSCAR-2109 881,722 OSCAR-2301 704,312 OSCAR-2019 703,991 OSCAR-2201 439,778 wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/romanian-corpus.text10M<n<100M1 likes1k downloads6mo agoHugging Face06RoMoDataset /RoMo-SMPL RoMo-SMPL — In-the-Wild SMPL Body Motion (RoMo Paper Core) RoMo-SMPL is the paper-aligned release of the RoMo body motion corpus in SMPL body parameter space (global orientation, 21-joint body pose, shape, translation). Each clip includes five text captions and a three-level semantic taxonomy (category, subcategory, atomic action), with fixed train / val / test splits. Paper: RoMo: A Large-Scale, Richly Organized Dataset and Semantic Taxonomy for Human Motion Generation… See the full description on the dataset page: https://huggingface.co/datasets/RoMoDataset/RoMo-SMPL.texttext-to-3d100K<n<1M0 likes894 downloads4mo agoHugging Face07rotarue /fineweb2-romanian-shardstext10M<n<100M1 likes639 downloads10mo agoHugging Face08romrawinjp /multilingual-coco Multilingual Common Objects in Context (COCO) Dataset This dataset is a collection of multiple language open-source captions of COCO dataset. The split in this dataset is set according to Andrej Karpathy's split from dataset_coco.json file. The collection was created specifically for simplicity of use in training and evaluation pipeline by non-commercial and research purposes. The COCO images dataset is licensed under a Creative Commons Attribution 4.0 License.… See the full description on the dataset page: https://huggingface.co/datasets/romrawinjp/multilingual-coco.imageimage-to-text100K<n<1M3 likes534 downloads2y agoHugging Face09Roman1111111 /claude-sonnet-4.6-120000xlicense: mit task_categories: text-generation text2text-generation language: en tags: reasoning uncensored math code claude-sonnet-4.6 claude-opus-4.6 gemini-3.1-pro size_categories: 100K<n<1M Please support if possible claude-sonnet-4.6-natural-large Sonnet4.6 NATURAL REASONING Multi-Domain(covered all possible topics in chats)/ Uncensored generated by claude sonnet 4.6(my biggest and most expensive project, i spent all my birthday money gifts for you guys❤️😁😭😭😭) 01… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/claude-sonnet-4.6-120000x.text100K<n<1M83 likes502 downloads5mo agoHugging Face10Sudehsna /Romansh_German_Parallel_Data Romansh–German Parallel Dataset (FineWeb-Based) This dataset contains automatically aligned Romansh–German document pairs, extracted from the Fineweb2 using cosine similarity over OpenAI embeddings. It was created as part of a university programming project focused on document-level parallel data extraction. Description This project performs document-level alignment between Romansh and German web texts, which were extracted from the Fineweb2 dataset. It uses OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/Sudehsna/Romansh_German_Parallel_Data.tabular10K<n<100K2 likes469 downloads1y agoHugging Face11EurekaTian /ROMA_proactive ROMA Proactive Streaming Dataset Figure: Overview of ROMA's Streaming Dataset. This repository contains the Proactive subset (Green and Purple sections). Dataset Summary This repository contains the Proactive Interaction subset of the dataset introduced in the paper ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding. This dataset is designed to train multimodal models for streaming video understanding, specifically focusing on tasks… See the full description on the dataset page: https://huggingface.co/datasets/EurekaTian/ROMA_proactive.textvisual-question-answering100K<n<1M1 likes452 downloads8mo agoHugging Face12agentlans /rombodawg-Everything_Instruct Everything-Instruct: Supervised Finetuning Dataset This dataset contains over 7 000 000 instruction-response pairs for supervised fine-tuning large language models. It combines the following datasets: rombodawg/Everything_Instruct rombodawg/Everything_Instruct_Multilingual It can be used for: Improving code generation and debugging Enhancing creative writing Improving general instruction followingFor English and many other languages Processing Removing duplicate… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/rombodawg-Everything_Instruct.texttext-generation1M<n<10M0 likes401 downloads9mo agoHugging Face13eduardem /romanian-speech-v2 Research Use Only — This dataset is released strictly for personal research and educational purposes. The processing pipeline and all scripts are fully open source, but the underlying audio originates from sources with varying copyrights. Only short fragments were used under fair use provisions and EU Copyright Directive Art. 3 (text and data mining for scientific research). This dataset must not be used for redistribution of the source material, commercial purposes, or training commercially… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/romanian-speech-v2.audiotext-to-speech100K<n<1M2 likes401 downloads7mo agoHugging Face14himalaya-ai /nepali-roman-pretraintext10M<n<100M0 likes388 downloads6mo agoHugging Face15BAAI /ROME 🏠Project Page & Leaderboard | 💻Code | 📄Paper | 🤗Data | 🤗Evaluation Response This repository contains a visual reasoning benchmark named ROME from the paper FlagEval Findings Report: A Preliminary Evaluation of Large Reasoning Models on Automatically Verifiable Textual and Visual Questions. ROME include 8 subtasks (281 high-quality questions in total). Each sample has been verified to ensure that images are necessary to answer correctly: Academic questions from college courses Diagrams… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/ROME.imageimage-text-to-textn<1K5 likes384 downloads1y agoHugging Face16Yuyeong /rw_roman-empire_standard_1_masktabular10M<n<100M0 likes383 downloads1y agoHugging Face17Yuyeong /rw_roman-empire_standard_2_masktabular10M<n<100M0 likes356 downloads1y agoHugging Face18Roman1111111 /gemini-3.1-pro-hard-high-reasoning Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M Dataset Details Dataset Description This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification. The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3.1-pro-hard-high-reasoning.textquestion-answering1K<n<10K62 likes351 downloads7mo agoHugging Face19romiroll /logical-reasoning-qa-dataset Dataset Card for "logical-reasoning-qa-dataset" More Information needed textn<1K0 likes350 downloads1y agoHugging Face20Yuyeong /rw_roman-empire_mdlr_6_masktabular10M<n<100M0 likes348 downloads1y agoHugging Face21Yuyeong /rw_roman-empire_node2vec_1_masktabular10M<n<100M0 likes345 downloads1y agoHugging Face22community-datasets /roman_urdu_hate_speech Dataset Card for roman_urdu_hate_speech Dataset Summary The Roman Urdu Hate-Speech and Offensive Language Detection (RUHSOLD) dataset is a Roman Urdu dataset of tweets annotated by experts in the relevant language. The authors develop the gold-standard for two sub-tasks. First sub-task is based on binary labels of Hate-Offensive content and Normal content (i.e., inoffensive language). These labels are self-explanatory. The authors refer to this sub-task as coarse-grained… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/roman_urdu_hate_speech.texttext-classification10K<n<100K3 likes330 downloads2y agoHugging Face23Yuyeong /rw_roman-empire_nbw_2_masktabular10M<n<100M0 likes330 downloads1y agoHugging Face24RoMoDataset /RoMo-HML-263 RoMo-HML-263 — RoMo Body Motion in HumanML3D-263 Features RoMo-HML-263 is the RoMo body corpus packed in the 263-dimensional HumanML3D motion-feature representation, paired with rich multi-level text descriptions. It is the drop-in companion for training and evaluating models built around the HumanML3D feature set, sized at the RoMo scale (~815K clips). ⚠️ Access: This dataset is currently private / internal. It will be released publicly in conjunction with the RoMo paper.… See the full description on the dataset page: https://huggingface.co/datasets/RoMoDataset/RoMo-HML-263.texttext-to-3d100K<n<1M0 likes330 downloads4mo agoHugging Face25Roman1111111 /opus-gpt-swe-frontier-core SWE Base Repository-level software engineering trajectories for training coding agents. 2,459 chat trajectories · 48,499 API calls · $837.57 recorded generation cost SWE-bench · debugging · patching · tools · agents Overview SWE Base is a software-engineering dataset centered on real repository issues. Each training example gives an agent a problem statement and captures the multi-turn process of inspecting a codebase, reasoning about a bug… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/opus-gpt-swe-frontier-core.tabulartext-generation1K<n<10K3 likes328 downloads1mo agoHugging Face26Yuyeong /rw_roman-empire_node2vec_6_masktabular10M<n<100M0 likes321 downloads1y agoHugging Face27Yuyeong /rw_roman-empire_nbw_1_masktabular10M<n<100M0 likes321 downloads1y agoHugging Face28Yuyeong /rw_roman-empire_nbw_6_masktabular10M<n<100M0 likes321 downloads1y agoHugging Face29Yuyeong /rw_roman-empire_mdlr_2_masktabular10M<n<100M0 likes318 downloads1y agoHugging Face30FraPiz /moldovan-dialectal-romanian-speech-corpus Moldovan Dialectal Romanian Educational Speech Corpus This dataset contains aligned Romanian educational speech with Moldovan dialectal characteristics. It was constructed from publicly accessible lesson videos recorded by teachers from the Republic of Moldova and published through the EducatieOnline platform. The corpus supports research on automatic speech recognition (ASR), text-to-speech synthesis (TTS), forced alignment, and low-resource dialectal speech processing.… See the full description on the dataset page: https://huggingface.co/datasets/FraPiz/moldovan-dialectal-romanian-speech-corpus.audioautomatic-speech-recognition10K<n<100K1 likes316 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.