datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KoHRM-Text-1.4B-sft-lora-data
KoHRM-Text-1.4B SFT and LoRA Prepared Data
This dataset repo stores curated KoHRM SFT/LoRA subsets in the same tokenized
HRM-Text V1Dataset format used by training. It is intended for quick behavior
alignment experiments after KoHRM pretraining.
Model repo:
https://huggingface.co/LLM-OS-Models/KoHRM-Text-1.4B
Code repo:
https://github.com/LLM-OS-Models/KoHRM-text
Format
Each folder is a prepared V1Dataset:
<dataset-name>/
metadata.json
tokenizer_info.json… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/KoHRM-Text-1.4B-sft-lora-data.Classical-Mechanics-Equations-Dataset_SFT-or-LoRA
Classical Mechanics Equations Dataset (SFT / LoRA Ready)
A structured dataset of 64 classical mechanics equations from Newtonian,
Lagrangian, and Hamiltonian mechanics, expanded into 448 instruction-tuning
rows across three task types: equation explanation, Q&A, and derivation.
Designed for fine-tuning LLMs on physics reasoning, STEM Q&A, and
equation understanding tasks.
Overview
Property
Value
Domain
Classical Mechanics (Physics)
Total rows
448
Train… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/Classical-Mechanics-Equations-Dataset_SFT-or-LoRA.lora-emotional-alignment-sample
BrightRun BrightRun Emotional Alignment Dataset — Sample Preview
🎯 Train Your LLM to Handle Emotionally Complex Conversations
This is a 12-conversation sample. The full dataset contains 242 conversations and 1,567 training pairs.
⚠️ This is a Sample — Not the Full Dataset
You're looking at 12 sample conversations designed to help you evaluate data quality before downloading the complete dataset.
What You Get Here
What You Get at brighthub.ai… See the full description on the dataset page: https://huggingface.co/datasets/BrightHubAI/lora-emotional-alignment-sample.Jazz-Blues-Music-Dataset_SFT-or-LoRA
Jazz & Blues Music Dataset (SFT / LoRA Ready)
A structured dataset covering 82 iconic Jazz and Blues songs, 21 artist
profiles, and 41 historical events, expanded into 1,219
instruction-tuning rows across 7 task types.
Designed for fine-tuning LLMs on music knowledge, cultural history, artist
biography, and domain-specific Q&A tasks.
Overview
Property
Value
Domain
Jazz & Blues Music
Total rows
1,219
Train split
1,036 (85%)
Validation split
91 (~7.5%)… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/Jazz-Blues-Music-Dataset_SFT-or-LoRA.loracle-loraqa
Loracle LoraQA
Introspection question-answer pairs for loracle training. Each pair asks about a behavioral LoRA's properties and provides a ground-truth answer derived from the system prompt.
Generation
Model: Gemini 3.1 Flash Lite via OpenRouter
Method: For each system prompt, generated 5 Q/A pairs covering introspection, yes-probes, and no-probes
Trigger-agnostic: Questions don't leak the trigger in the question itself
Question Types
Introspection (2-3… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/loracle-loraqa.loracle-ia-diverse-qa
loracle-ia-diverse-qa — v6
QA training data for the loracle — a model that reads a LoRA's weight deltas and answers questions about the behavior it encodes.
What this is
Each row pairs a LoRA identifier with a (question, answer) where the answer requires reading the LoRA's direction-token projections to answer correctly. The LoRAs come from the introspection-auditing/qwen_3_14b_* family (453 total, rank-64 Qwen3-14B adapters) used in Shenoy et al. (2026) Introspection… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/loracle-ia-diverse-qa.LoRA_Fine-Tune_Q_and_A
Introduction
This dataset is used to fine-tune Deepseek-R1.
