CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01EurekaTian /ROMA_proactive ROMA Proactive Streaming Dataset Figure: Overview of ROMA's Streaming Dataset. This repository contains the Proactive subset (Green and Purple sections). Dataset Summary This repository contains the Proactive Interaction subset of the dataset introduced in the paper ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding. This dataset is designed to train multimodal models for streaming video understanding, specifically focusing on tasks… See the full description on the dataset page: https://huggingface.co/datasets/EurekaTian/ROMA_proactive.textvisual-question-answering100K<n<1M1 likes519 downloads8mo agoHugging Face02agentlans /rombodawg-Everything_Instruct Everything-Instruct: Supervised Finetuning Dataset This dataset contains over 7 000 000 instruction-response pairs for supervised fine-tuning large language models. It combines the following datasets: rombodawg/Everything_Instruct rombodawg/Everything_Instruct_Multilingual It can be used for: Improving code generation and debugging Enhancing creative writing Improving general instruction followingFor English and many other languages Processing Removing duplicate… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/rombodawg-Everything_Instruct.texttext-generation1M<n<10M0 likes384 downloads9mo agoHugging Face03Roman1111111 /gemini-3.1-pro-hard-high-reasoning Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M Dataset Details Dataset Description This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification. The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3.1-pro-hard-high-reasoning.textquestion-answering1K<n<10K62 likes336 downloads7mo agoHugging Face04Roman1111111 /opus-gpt-swe-frontier-core SWE Base Repository-level software engineering trajectories for training coding agents. 2,459 chat trajectories · 48,499 API calls · $837.57 recorded generation cost SWE-bench · debugging · patching · tools · agents Overview SWE Base is a software-engineering dataset centered on real repository issues. Each training example gives an agent a problem statement and captures the multi-turn process of inspecting a codebase, reasoning about a bug… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/opus-gpt-swe-frontier-core.tabulartext-generation1K<n<10K3 likes298 downloads1mo agoHugging Face05ajaysinghsat /aosp-rom-dataset Modern OnePlus & Xiaomi Flagship AOSP Engineering Dataset Structured, lightweight multi-domain shards covering: device_trees_kernel: OnePlus (SM8550/SM8450), Xiaomi (SM8550/SM8450), GKI modules, BoardConfig, and DTS bindings. systemui_launcher: Lawnchair spring physics, AGSL runtime shaders, and Avium/Axion/PixelOS UI panels. build_sepolicy: Soong blueprints (Android.bp), product makefiles, and strict SELinux rules (.te). frameworks: Native SurfaceFlinger C++ compositor and… See the full description on the dataset page: https://huggingface.co/datasets/ajaysinghsat/aosp-rom-dataset.texttext-generation100K<n<1M0 likes133 downloads23d agoHugging Face06Roman1111111 /gemini-3-pro-10000x-hard-high-reasoning Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning Dataset Details Dataset Description Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement. This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3-pro-10000x-hard-high-reasoning.textquestion-answering10K<n<100K57 likes92 downloads7mo agoHugging Face07Ashar086 /roman-urdu-qwen25-3b-blindspot Roman Urdu / code-switch blind spot (Qwen2.5-3B-Instruct) Hand-built eval: 8 Pakistani situations x 3 surfaces (English, formal Urdu, Roman Urdu). Model: Qwen/Qwen2.5-3B-Instruct. Greedy decoding, 4-bit, Colab T4. Evaluation condition pass n rate english 7 8 0.88 formal_urdu 3 8 0.38 roman_urdu 1 8 0.12 Files: prompts.jsonl, outputs.jsonl, judged.jsonl, scores.json Roman Urdu traces: 01_ro: NADRA described as a motor-vehicle department 02_ro:… See the full description on the dataset page: https://huggingface.co/datasets/Ashar086/roman-urdu-qwen25-3b-blindspot.texttext-generationn<1K0 likes45 downloads4d agoHugging Face08Redgerd /roman-urdu-alpaca-qa-mix Dataset Card for Roman Urdu + Alpaca QA Mix This dataset is intended to support fine-tuning and evaluation of language models that understand and respond to Roman Urdu and English instructions. It consists of 1,022 records in total: 500 examples in Roman Urdu generated from high-quality Urdu sources and transliterated using the ChatGPT API. 500 examples in English randomly sampled from the Stanford Alpaca dataset. The dataset follows the same format as Alpaca-style instruction… See the full description on the dataset page: https://huggingface.co/datasets/Redgerd/roman-urdu-alpaca-qa-mix.textquestion-answering1K<n<10K0 likes35 downloads1y agoHugging Face09wasabiP /japanese-triplet-lifestyle-romance 🏯 Japanese Preference Dataset: Counseling & Advice (Free Sample) This repository provides a free sample of a Japanese preference learning dataset designed for Direct Preference Optimization (DPO), RLHF, Reward Modeling, response ranking, and Japanese LLM alignment. The dataset focuses on realistic Japanese counseling and advice scenarios, helping language models learn not only factual correctness but also empathy, contextual understanding, and practical response quality.… See the full description on the dataset page: https://huggingface.co/datasets/wasabiP/japanese-triplet-lifestyle-romance.texttext-generationn<1K0 likes33 downloads2mo agoHugging Face10benjleite /FairytaleQA-translated-romanian Dataset Card for FairytaleQA-translated-ptBR Dataset Summary This repository contains the Romanian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-romanian.textquestion-answering10K<n<100K1 likes31 downloads1y agoHugging Face11Roman190928 /Wikitext I made this by downloading and extracting the wikidump (smaller version) This might be better for fine tuning its a jsonl file if you're low on space, ill be giving the zipped jsonl too :D 385,692 lines or around 385,000 texttext-generation100K<n<1M0 likes28 downloads11mo agoHugging Face12Ilia-Iliev /romani_compasito Romani Prompts (Kompasito) Prompt/context pairs derived from Kompasito — Manual on Human Rights Education for Children (Council of Europe), translated into Romani. Built for adapting LLMs to Romani via SFT/DPO. Structure Each row is a JSON object: prompt — an English instruction sampled from one of three buckets: comprehension, translation, or generative. context — a Romani text chunk (~1000 chars, 150-char overlap) from the source PDF. Each surviving chunk appears in 3… See the full description on the dataset page: https://huggingface.co/datasets/Ilia-Iliev/romani_compasito.texttext-generation1K<n<10K0 likes20 downloads5mo agoHugging Face13xaviviro /FEDERICO-GARCIA-LORCA-canciones-poemas-romances-annotated Federico García Lorca - Annotated Poetry Dataset A curated and annotated dataset of 283 poems by Federico García Lorca, spanning 9 of his major works (1921--1940). Each poem is enriched with publication metadata and GPT-4-generated thematic and contextual annotations. Use Case: LLM Generalization Evaluation This dataset was created to evaluate how well large language models can generalize literary style from a small, domain-specific corpus. It has been used to fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/xaviviro/FEDERICO-GARCIA-LORCA-canciones-poemas-romances-annotated.tabulartext-generationn<1K0 likes16 downloads7mo agoHugging Face14RomainPct /steve-jobs-question-and-answerstexttext-generationn<1K0 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.