datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ROMA_proactive
ROMA Proactive Streaming Dataset
Figure: Overview of ROMA's Streaming Dataset. This repository contains the Proactive subset (Green and Purple sections).
Dataset Summary
This repository contains the Proactive Interaction subset of the dataset introduced in the paper ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding.
This dataset is designed to train multimodal models for streaming video understanding, specifically focusing on tasks… See the full description on the dataset page: https://huggingface.co/datasets/EurekaTian/ROMA_proactive.rombodawg-Everything_Instruct
Everything-Instruct: Supervised Finetuning Dataset
This dataset contains over 7 000 000 instruction-response pairs for supervised fine-tuning large language models.
It combines the following datasets:
rombodawg/Everything_Instruct
rombodawg/Everything_Instruct_Multilingual
It can be used for:
Improving code generation and debugging
Enhancing creative writing
Improving general instruction followingFor English and many other languages
Processing
Removing duplicate… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/rombodawg-Everything_Instruct.gemini-3.1-pro-hard-high-reasoning
Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M
Dataset Details
Dataset Description
This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification.
The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3.1-pro-hard-high-reasoning.opus-gpt-swe-frontier-core
SWE Base
Repository-level software engineering trajectories for training coding agents.
2,459 chat trajectories · 48,499 API calls · $837.57 recorded generation cost
SWE-bench · debugging · patching · tools · agents
Overview
SWE Base is a software-engineering dataset centered on real repository issues. Each training example gives an agent a problem statement and captures the multi-turn process of inspecting a codebase, reasoning about a bug… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/opus-gpt-swe-frontier-core.aosp-rom-dataset
Modern OnePlus & Xiaomi Flagship AOSP Engineering Dataset
Structured, lightweight multi-domain shards covering:
device_trees_kernel: OnePlus (SM8550/SM8450), Xiaomi (SM8550/SM8450), GKI modules, BoardConfig, and DTS bindings.
systemui_launcher: Lawnchair spring physics, AGSL runtime shaders, and Avium/Axion/PixelOS UI panels.
build_sepolicy: Soong blueprints (Android.bp), product makefiles, and strict SELinux rules (.te).
frameworks: Native SurfaceFlinger C++ compositor and… See the full description on the dataset page: https://huggingface.co/datasets/ajaysinghsat/aosp-rom-dataset.gemini-3-pro-10000x-hard-high-reasoning
Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning
Dataset Details
Dataset Description
Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement.
This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3-pro-10000x-hard-high-reasoning.roman-urdu-qwen25-3b-blindspot
Roman Urdu / code-switch blind spot (Qwen2.5-3B-Instruct)
Hand-built eval: 8 Pakistani situations x 3 surfaces (English, formal Urdu, Roman Urdu).
Model: Qwen/Qwen2.5-3B-Instruct. Greedy decoding, 4-bit, Colab T4.
Evaluation
condition
pass
n
rate
english
7
8
0.88
formal_urdu
3
8
0.38
roman_urdu
1
8
0.12
Files: prompts.jsonl, outputs.jsonl, judged.jsonl, scores.json
Roman Urdu traces:
01_ro: NADRA described as a motor-vehicle department
02_ro:… See the full description on the dataset page: https://huggingface.co/datasets/Ashar086/roman-urdu-qwen25-3b-blindspot.roman-urdu-alpaca-qa-mix
Dataset Card for Roman Urdu + Alpaca QA Mix
This dataset is intended to support fine-tuning and evaluation of language models that understand and respond to Roman Urdu and English instructions. It consists of 1,022 records in total:
500 examples in Roman Urdu generated from high-quality Urdu sources and transliterated using the ChatGPT API.
500 examples in English randomly sampled from the Stanford Alpaca dataset.
The dataset follows the same format as Alpaca-style instruction… See the full description on the dataset page: https://huggingface.co/datasets/Redgerd/roman-urdu-alpaca-qa-mix.japanese-triplet-lifestyle-romance
🏯 Japanese Preference Dataset: Counseling & Advice (Free Sample)
This repository provides a free sample of a Japanese preference learning dataset designed for Direct Preference Optimization (DPO), RLHF, Reward Modeling, response ranking, and Japanese LLM alignment.
The dataset focuses on realistic Japanese counseling and advice scenarios, helping language models learn not only factual correctness but also empathy, contextual understanding, and practical response quality.… See the full description on the dataset page: https://huggingface.co/datasets/wasabiP/japanese-triplet-lifestyle-romance.FairytaleQA-translated-romanian
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Romanian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-romanian.Wikitext
I made this by downloading and extracting the wikidump (smaller version)
This might be better for fine tuning
its a jsonl file
if you're low on space, ill be giving the zipped jsonl too :D
385,692 lines
or around 385,000
romani_compasito
Romani Prompts (Kompasito)
Prompt/context pairs derived from Kompasito — Manual on Human Rights Education for Children (Council of Europe), translated into Romani. Built for adapting LLMs to Romani via SFT/DPO.
Structure
Each row is a JSON object:
prompt — an English instruction sampled from one of three buckets: comprehension, translation, or generative.
context — a Romani text chunk (~1000 chars, 150-char overlap) from the source PDF.
Each surviving chunk appears in 3… See the full description on the dataset page: https://huggingface.co/datasets/Ilia-Iliev/romani_compasito.FEDERICO-GARCIA-LORCA-canciones-poemas-romances-annotated
Federico García Lorca - Annotated Poetry Dataset
A curated and annotated dataset of 283 poems by Federico García Lorca, spanning 9 of his major works (1921--1940). Each poem is enriched with publication metadata and GPT-4-generated thematic and contextual annotations.
Use Case: LLM Generalization Evaluation
This dataset was created to evaluate how well large language models can generalize literary style from a small, domain-specific corpus. It has been used to fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/xaviviro/FEDERICO-GARCIA-LORCA-canciones-poemas-romances-annotated.steve-jobs-question-and-answers
