CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RoMoDataset /RoMo-SMPL RoMo-SMPL — In-the-Wild SMPL Body Motion (RoMo Paper Core) RoMo-SMPL is the paper-aligned release of the RoMo body motion corpus in SMPL body parameter space (global orientation, 21-joint body pose, shape, translation). Each clip includes five text captions and a three-level semantic taxonomy (category, subcategory, atomic action), with fixed train / val / test splits. Paper: RoMo: A Large-Scale, Richly Organized Dataset and Semantic Taxonomy for Human Motion Generation… See the full description on the dataset page: https://huggingface.co/datasets/RoMoDataset/RoMo-SMPL.texttext-to-3d100K<n<1M0 likes901 downloads4mo agoHugging Face02EurekaTian /ROMA_proactive ROMA Proactive Streaming Dataset Figure: Overview of ROMA's Streaming Dataset. This repository contains the Proactive subset (Green and Purple sections). Dataset Summary This repository contains the Proactive Interaction subset of the dataset introduced in the paper ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding. This dataset is designed to train multimodal models for streaming video understanding, specifically focusing on tasks… See the full description on the dataset page: https://huggingface.co/datasets/EurekaTian/ROMA_proactive.textvisual-question-answering100K<n<1M1 likes519 downloads8mo agoHugging Face03agentlans /rombodawg-Everything_Instruct Everything-Instruct: Supervised Finetuning Dataset This dataset contains over 7 000 000 instruction-response pairs for supervised fine-tuning large language models. It combines the following datasets: rombodawg/Everything_Instruct rombodawg/Everything_Instruct_Multilingual It can be used for: Improving code generation and debugging Enhancing creative writing Improving general instruction followingFor English and many other languages Processing Removing duplicate… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/rombodawg-Everything_Instruct.texttext-generation1M<n<10M0 likes384 downloads9mo agoHugging Face04RoMoDataset /RoMo-HML-263 RoMo-HML-263 — RoMo Body Motion in HumanML3D-263 Features RoMo-HML-263 is the RoMo body corpus packed in the 263-dimensional HumanML3D motion-feature representation, paired with rich multi-level text descriptions. It is the drop-in companion for training and evaluating models built around the HumanML3D feature set, sized at the RoMo scale (~815K clips). ⚠️ Access: This dataset is currently private / internal. It will be released publicly in conjunction with the RoMo paper.… See the full description on the dataset page: https://huggingface.co/datasets/RoMoDataset/RoMo-HML-263.texttext-to-3d100K<n<1M0 likes338 downloads4mo agoHugging Face05Roman1111111 /gemini-3.1-pro-hard-high-reasoning Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M Dataset Details Dataset Description This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification. The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3.1-pro-hard-high-reasoning.textquestion-answering1K<n<10K62 likes336 downloads7mo agoHugging Face06Roman1111111 /opus-gpt-swe-frontier-core SWE Base Repository-level software engineering trajectories for training coding agents. 2,459 chat trajectories · 48,499 API calls · $837.57 recorded generation cost SWE-bench · debugging · patching · tools · agents Overview SWE Base is a software-engineering dataset centered on real repository issues. Each training example gives an agent a problem statement and captures the multi-turn process of inspecting a codebase, reasoning about a bug… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/opus-gpt-swe-frontier-core.tabulartext-generation1K<n<10K3 likes298 downloads1mo agoHugging Face07RoMoDataset /RoMo-272 RoMo-272 — RoMo Body Motion in the 272-D Motion Representation RoMo-272 is the RoMo body-motion corpus (paper core) packed in the 272-dimensional motion representation of Li-xingXiao/272-dim-Motion-Representation — a HumanML3D-derived encoding that augments the standard 263-D HumanML3D layout with 9 additional absolute/global features used by several recent text-to-motion methods. Each clip carries five text captions and a three-level semantic taxonomy, with fixed train / val /… See the full description on the dataset page: https://huggingface.co/datasets/RoMoDataset/RoMo-272.texttext-to-3d100K<n<1M0 likes227 downloads4mo agoHugging Face08RoMoDataset /RoMo-SOMA-77 RoMo-SOMA-77 — RoMo Body+Hand Motion in 933-D Kimodo SOMA-77 Features RoMo-SOMA-77 is the RoMo body+hand corpus packed in a 933-dimensional Kimodo SOMA-77 motion-feature representation, paired with rich multi-level text descriptions. It is the publication target for the SOMA-based body-and-hand model family. Scope: paper-core (romo_official = True), matching RoMo-SMPL, RoMo-HML-263, and RoMo-272. A small number of clips are dropped where SOMA conversion produced non-finite… See the full description on the dataset page: https://huggingface.co/datasets/RoMoDataset/RoMo-SOMA-77.texttext-to-3d100K<n<1M0 likes225 downloads4mo agoHugging Face09d0rj /ROMB-1.0 ♦ ROMB Русское описание и инструкция ROMB (Russian Olympiad Math Benchmark) evaluates models on Russian-language school olympiad mathematics. The test set contains 2552 text-only tasks: 1716 arithmetic/other tasks, 644 logic tasks, and 192 geometry tasks. Tasks have typed answers, answer-format notes, and per-task checking rules. The evaluator also supports configurable v3 runs: native thinking, optional JSON Schema constrained decoding, plain or \boxed{…} answers, and… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/ROMB-1.0.texttext-generation1K<n<10K1 likes142 downloads16d agoHugging Face10ajaysinghsat /aosp-rom-dataset Modern OnePlus & Xiaomi Flagship AOSP Engineering Dataset Structured, lightweight multi-domain shards covering: device_trees_kernel: OnePlus (SM8550/SM8450), Xiaomi (SM8550/SM8450), GKI modules, BoardConfig, and DTS bindings. systemui_launcher: Lawnchair spring physics, AGSL runtime shaders, and Avium/Axion/PixelOS UI panels. build_sepolicy: Soong blueprints (Android.bp), product makefiles, and strict SELinux rules (.te). frameworks: Native SurfaceFlinger C++ compositor and… See the full description on the dataset page: https://huggingface.co/datasets/ajaysinghsat/aosp-rom-dataset.texttext-generation100K<n<1M0 likes133 downloads23d agoHugging Face11RomanCast /WikiSpell_custom WikiSpell Description This dataset is a custom implementation of the WikiSpell dataset introduced in Character-Aware Models Improve Visual Text Rendering by Liu et al. (2022). Similarly to the original WikiSpell dataset, the training set is composed of 5000 words taken uniformly from the 50% least common Wiktionary words (taken from this Wiktionary extraction), and 5000 words sampled according to their frequencies taken from the 50% most common Wiktionary words. The… See the full description on the dataset page: https://huggingface.co/datasets/RomanCast/WikiSpell_custom.texttext-generation10K<n<100K0 likes115 downloads3y agoHugging Face12Roman1111111 /gemini-3-pro-10000x-hard-high-reasoning Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning Dataset Details Dataset Description Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement. This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3-pro-10000x-hard-high-reasoning.textquestion-answering10K<n<100K57 likes92 downloads7mo agoHugging Face13romban38 /reddit_dataset_51 Bittensor Subnet 13 Reddit Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more… See the full description on the dataset page: https://huggingface.co/datasets/romban38/reddit_dataset_51.texttext-classification1M<n<10M0 likes79 downloads1y agoHugging Face14louisbrulenaudet /Romulus-cpt-fr Romulus, continually pre-trained models for French law. Romulus is a series of continually pre-trained models enriched in French law and intended to serve as the basis for a fine-tuning process on labeled data. Please note that these models have not been aligned for the production of usable text as they stand, and will certainly need to be fine-tuned for the desired tasks in order to produce satisfactory results. The training corpus is made up of around 34,864,949 tokens… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/Romulus-cpt-fr.tabulartext-generation100K<n<1M5 likes74 downloads2y agoHugging Face15romban38 /x_dataset_51 Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/romban38/x_dataset_51.texttext-classification100M<n<1B0 likes62 downloads1y agoHugging Face16Ashar086 /roman-urdu-qwen25-3b-blindspot Roman Urdu / code-switch blind spot (Qwen2.5-3B-Instruct) Hand-built eval: 8 Pakistani situations x 3 surfaces (English, formal Urdu, Roman Urdu). Model: Qwen/Qwen2.5-3B-Instruct. Greedy decoding, 4-bit, Colab T4. Evaluation condition pass n rate english 7 8 0.88 formal_urdu 3 8 0.38 roman_urdu 1 8 0.12 Files: prompts.jsonl, outputs.jsonl, judged.jsonl, scores.json Roman Urdu traces: 01_ro: NADRA described as a motor-vehicle department 02_ro:… See the full description on the dataset page: https://huggingface.co/datasets/Ashar086/roman-urdu-qwen25-3b-blindspot.texttext-generationn<1K0 likes45 downloads3d agoHugging Face17Redgerd /roman-urdu-alpaca-qa-mix Dataset Card for Roman Urdu + Alpaca QA Mix This dataset is intended to support fine-tuning and evaluation of language models that understand and respond to Roman Urdu and English instructions. It consists of 1,022 records in total: 500 examples in Roman Urdu generated from high-quality Urdu sources and transliterated using the ChatGPT API. 500 examples in English randomly sampled from the Stanford Alpaca dataset. The dataset follows the same format as Alpaca-style instruction… See the full description on the dataset page: https://huggingface.co/datasets/Redgerd/roman-urdu-alpaca-qa-mix.textquestion-answering1K<n<10K0 likes35 downloads1y agoHugging Face18wasabiP /japanese-triplet-lifestyle-romance 🏯 Japanese Preference Dataset: Counseling & Advice (Free Sample) This repository provides a free sample of a Japanese preference learning dataset designed for Direct Preference Optimization (DPO), RLHF, Reward Modeling, response ranking, and Japanese LLM alignment. The dataset focuses on realistic Japanese counseling and advice scenarios, helping language models learn not only factual correctness but also empathy, contextual understanding, and practical response quality.… See the full description on the dataset page: https://huggingface.co/datasets/wasabiP/japanese-triplet-lifestyle-romance.texttext-generationn<1K0 likes33 downloads2mo agoHugging Face19benjleite /FairytaleQA-translated-romanian Dataset Card for FairytaleQA-translated-ptBR Dataset Summary This repository contains the Romanian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-romanian.textquestion-answering10K<n<100K1 likes31 downloads1y agoHugging Face20pythainlp /thai-romanization-datasetDataset from github.com/wannaphong/thai-romanization/blob/master/dataset/data.csv texttext-generation100K<n<1M0 likes29 downloads1y agoHugging Face21Roman190928 /Wikitext I made this by downloading and extracting the wikidump (smaller version) This might be better for fine tuning its a jsonl file if you're low on space, ill be giving the zipped jsonl too :D 385,692 lines or around 385,000 texttext-generation100K<n<1M0 likes28 downloads11mo agoHugging Face22nirajandhakal /Devnagari-Romanized-Pair Dataset Overview The dataset devanagari romanized pair contains, 959 rows, where each row has one English sentence and its corresponding Nepali translations both in the devanagari script and in romanized format, the size of the data set is less than 1,000 elements and it's designed for use in Translation, text generation and text to text generation tasks. texttranslationn<1K3 likes26 downloads2y agoHugging Face23Yxanul /English-Romanian-Magpie-Reasoning English-Romanian Translation Pairs from Magpie-Reasoning This dataset contains 150,000 high-quality English-Romanian parallel translation pairs derived from the Magpie-Reasoning dataset, specifically designed for training and evaluating machine translation models with a focus on technical, mathematical, and code-related content. Source Datasets This dataset is created by aligning: English: Magpie-Align/Magpie-Reasoning-V1-150K Romanian: OpenLLM-Ro/ro_sft_magpie_reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/English-Romanian-Magpie-Reasoning.texttranslation100K<n<1M0 likes24 downloads11mo agoHugging Face24Ilia-Iliev /romani_compasito Romani Prompts (Kompasito) Prompt/context pairs derived from Kompasito — Manual on Human Rights Education for Children (Council of Europe), translated into Romani. Built for adapting LLMs to Romani via SFT/DPO. Structure Each row is a JSON object: prompt — an English instruction sampled from one of three buckets: comprehension, translation, or generative. context — a Romani text chunk (~1000 chars, 150-char overlap) from the source PDF. Each surviving chunk appears in 3… See the full description on the dataset page: https://huggingface.co/datasets/Ilia-Iliev/romani_compasito.texttext-generation1K<n<10K0 likes20 downloads5mo agoHugging Face25cosmadrian /romathgated RoMath - A Mathematical Reasoning Benchmarking Suite from Descriptions in 🇷🇴 Romanian 🇷🇴" This is the official dataset for RoMath. It currently is comprised of the test and training splits for RoMath-Baccalaureate, RoMath-Competitions and RoMath-Synthetic. Mathematics has long been conveyed through natural language, primarily for human understanding. With the rise of mechanized mathematics and proof assistants, there is a growing need to understand informal mathematical text… See the full description on the dataset page: https://huggingface.co/datasets/cosmadrian/romath.texttext-generation10K<n<100K3 likes17 downloads2y agoHugging Face26xaviviro /FEDERICO-GARCIA-LORCA-canciones-poemas-romances-annotated Federico García Lorca - Annotated Poetry Dataset A curated and annotated dataset of 283 poems by Federico García Lorca, spanning 9 of his major works (1921--1940). Each poem is enriched with publication metadata and GPT-4-generated thematic and contextual annotations. Use Case: LLM Generalization Evaluation This dataset was created to evaluate how well large language models can generalize literary style from a small, domain-specific corpus. It has been used to fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/xaviviro/FEDERICO-GARCIA-LORCA-canciones-poemas-romances-annotated.tabulartext-generationn<1K0 likes16 downloads7mo agoHugging Face27romizone /padel-indonesia-dataset 🏸 Padel Indonesia Dataset Dataset instruksi Q&A komprehensif tentang olahraga Padel di Indonesia Oleh Jekardah AI Labs — Romi Nur Ismanto Overview Dataset ini berisi 1,291 pasangan instruksi-output dalam Bahasa Indonesia yang mencakup seluruh aspek olahraga Padel di Indonesia. Dirancang untuk fine-tuning LLM agar bisa menjawab pertanyaan seputar Padel dengan akurat dan komprehensif. Format Setiap entry mengikuti format Alpaca-style: { "instruction":… See the full description on the dataset page: https://huggingface.co/datasets/romizone/padel-indonesia-dataset.textquestion-answering1K<n<10K0 likes12 downloads6mo agoHugging Face28ZenithVortex /romeo_and_juliet Dataset Card for romeo_and_juliet This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/ZenithVortex/romeo_and_juliet/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/ZenithVortex/romeo_and_juliet.texttext-generationn<1K0 likes8 downloads2y agoHugging Face29RomainPct /steve-jobs-question-and-answerstexttext-generationn<1K0 likes7 downloads2y agoHugging Face30RomeKim /my-distiset-3276ad2f Dataset Card for my-distiset-3276ad2f This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/RomeKim/my-distiset-3276ad2f/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/RomeKim/my-distiset-3276ad2f.texttext-generationn<1K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.