datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BOT_JORDANS-storageMassSpecGym
MassSpecGym provides a dataset and benchmark for the discovery and identification of new molecules from tandem mass spectrometry (MS/MS) spectra. The provided challenges abstract the process of scientific discovery of new molecules from biological and environmental samples into well-defined machine learning problems.
Papers
MassSpecGym in the Wild: Uncovering and Correcting Evaluation Pitfalls in AI-Driven Molecule Discovery (2025): Paper Link
MassSpecGym: A benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/roman-bushuiev/MassSpecGym.SWE-Bench-MultilingualC_CPPFileteredSWE-Bench-MultilingualC_CPPFiletered_newromanian-corpus
Romanian Text Corpus
A comprehensive, high-quality Romanian text corpus for language model pretraining.
Built by collecting and cleaning text from five Romanian-language sources.
Dataset Summary
Total documents: 19,886,412
Estimated tokens: ~20.8B
Language: Romanian (ro)
Format: Parquet (zstd compressed)
Source Breakdown
Source
Documents
mC4
16,875,310
OSCAR-2109
881,722
OSCAR-2301
704,312
OSCAR-2019
703,991
OSCAR-2201
439,778
wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/romanian-corpus.RoMo-SMPL
RoMo-SMPL — In-the-Wild SMPL Body Motion (RoMo Paper Core)
RoMo-SMPL is the paper-aligned release of the RoMo body motion corpus in SMPL body parameter space (global orientation, 21-joint body pose, shape, translation). Each clip includes five text captions and a three-level semantic taxonomy (category, subcategory, atomic action), with fixed train / val / test splits.
Paper: RoMo: A Large-Scale, Richly Organized Dataset and Semantic Taxonomy for Human Motion Generation… See the full description on the dataset page: https://huggingface.co/datasets/RoMoDataset/RoMo-SMPL.fineweb2-romanian-shardsmultilingual-coco
Multilingual Common Objects in Context (COCO) Dataset
This dataset is a collection of multiple language open-source captions of COCO dataset.
The split in this dataset is set according to Andrej Karpathy's split from dataset_coco.json file. The collection was created specifically for simplicity of use in training and evaluation pipeline by non-commercial and research purposes. The COCO images dataset is licensed under a Creative Commons Attribution 4.0 License.… See the full description on the dataset page: https://huggingface.co/datasets/romrawinjp/multilingual-coco.claude-sonnet-4.6-120000xlicense: mit
task_categories:
text-generation
text2text-generation
language:
en
tags:
reasoning
uncensored
math
code
claude-sonnet-4.6
claude-opus-4.6
gemini-3.1-pro
size_categories:
100K<n<1M
Please support if possible
claude-sonnet-4.6-natural-large
Sonnet4.6 NATURAL REASONING
Multi-Domain(covered all possible topics in chats)/ Uncensored generated by claude sonnet 4.6(my biggest and most expensive project, i spent all my birthday money gifts for you guys❤️😁😭😭😭)
01… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/claude-sonnet-4.6-120000x.Romansh_German_Parallel_Data
Romansh–German Parallel Dataset (FineWeb-Based)
This dataset contains automatically aligned Romansh–German document pairs, extracted from the Fineweb2 using cosine similarity over OpenAI embeddings. It was created as part of a university programming project focused on document-level parallel data extraction.
Description
This project performs document-level alignment between Romansh and German web texts, which were extracted from the Fineweb2 dataset. It uses OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/Sudehsna/Romansh_German_Parallel_Data.ROMA_proactive
ROMA Proactive Streaming Dataset
Figure: Overview of ROMA's Streaming Dataset. This repository contains the Proactive subset (Green and Purple sections).
Dataset Summary
This repository contains the Proactive Interaction subset of the dataset introduced in the paper ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding.
This dataset is designed to train multimodal models for streaming video understanding, specifically focusing on tasks… See the full description on the dataset page: https://huggingface.co/datasets/EurekaTian/ROMA_proactive.rombodawg-Everything_Instruct
Everything-Instruct: Supervised Finetuning Dataset
This dataset contains over 7 000 000 instruction-response pairs for supervised fine-tuning large language models.
It combines the following datasets:
rombodawg/Everything_Instruct
rombodawg/Everything_Instruct_Multilingual
It can be used for:
Improving code generation and debugging
Enhancing creative writing
Improving general instruction followingFor English and many other languages
Processing
Removing duplicate… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/rombodawg-Everything_Instruct.romanian-speech-v2
Research Use Only — This dataset is released strictly for personal research and educational
purposes. The processing pipeline and all scripts are fully open source, but the underlying audio
originates from sources with varying copyrights. Only short fragments were used under fair use
provisions and EU Copyright Directive Art. 3 (text and data mining for scientific research).
This dataset must not be used for redistribution of the source material, commercial purposes,
or training commercially… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/romanian-speech-v2.nepali-roman-pretrainROME
🏠Project Page & Leaderboard | 💻Code | 📄Paper | 🤗Data | 🤗Evaluation Response
This repository contains a visual reasoning benchmark named ROME from the paper FlagEval Findings Report: A Preliminary Evaluation of Large Reasoning Models on Automatically Verifiable Textual and Visual Questions.
ROME include 8 subtasks (281 high-quality questions in total). Each sample has been verified to ensure that images are necessary to answer correctly:
Academic
questions from college courses
Diagrams… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/ROME.rw_roman-empire_standard_1_maskrw_roman-empire_standard_2_maskgemini-3.1-pro-hard-high-reasoning
Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M
Dataset Details
Dataset Description
This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification.
The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3.1-pro-hard-high-reasoning.logical-reasoning-qa-dataset
Dataset Card for "logical-reasoning-qa-dataset"
More Information needed
rw_roman-empire_mdlr_6_maskrw_roman-empire_node2vec_1_maskroman_urdu_hate_speech
Dataset Card for roman_urdu_hate_speech
Dataset Summary
The Roman Urdu Hate-Speech and Offensive Language Detection (RUHSOLD) dataset is a Roman Urdu dataset of tweets annotated by experts in the relevant language. The authors develop the gold-standard for two sub-tasks. First sub-task is based on binary labels of Hate-Offensive content and Normal content (i.e., inoffensive language). These labels are self-explanatory. The authors refer to this sub-task as coarse-grained… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/roman_urdu_hate_speech.rw_roman-empire_nbw_2_maskRoMo-HML-263
RoMo-HML-263 — RoMo Body Motion in HumanML3D-263 Features
RoMo-HML-263 is the RoMo body corpus packed in the 263-dimensional HumanML3D motion-feature representation, paired with rich multi-level text descriptions. It is the drop-in companion for training and evaluating models built around the HumanML3D feature set, sized at the RoMo scale (~815K clips).
⚠️ Access: This dataset is currently private / internal. It will be released publicly in conjunction with the RoMo paper.… See the full description on the dataset page: https://huggingface.co/datasets/RoMoDataset/RoMo-HML-263.opus-gpt-swe-frontier-core
SWE Base
Repository-level software engineering trajectories for training coding agents.
2,459 chat trajectories · 48,499 API calls · $837.57 recorded generation cost
SWE-bench · debugging · patching · tools · agents
Overview
SWE Base is a software-engineering dataset centered on real repository issues. Each training example gives an agent a problem statement and captures the multi-turn process of inspecting a codebase, reasoning about a bug… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/opus-gpt-swe-frontier-core.rw_roman-empire_node2vec_6_maskrw_roman-empire_nbw_1_maskrw_roman-empire_nbw_6_maskrw_roman-empire_mdlr_2_maskmoldovan-dialectal-romanian-speech-corpus
Moldovan Dialectal Romanian Educational Speech Corpus
This dataset contains aligned Romanian educational speech with Moldovan
dialectal characteristics. It was constructed from publicly accessible lesson
videos recorded by teachers from the Republic of Moldova and published through
the EducatieOnline platform.
The corpus supports research on automatic speech recognition (ASR),
text-to-speech synthesis (TTS), forced alignment, and low-resource dialectal
speech processing.… See the full description on the dataset page: https://huggingface.co/datasets/FraPiz/moldovan-dialectal-romanian-speech-corpus.
