datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bio-safety-peft-lora
CBRN Safety Alignment & PEFT-LoRA Fine-Tuning Dataset
This repository contains the synthetic instruction-tuning dataset (.jsonl) designed for parameter-efficient fine-tuning (PEFT-LoRA) of edge language models (specifically Qwen/Qwen2.5-1.5B-Instruct).
The dataset is curated to evaluate and modify model logit distributions, persona attributions, and dual-use safety boundaries regarding Chemical, Biological, Radiological, and Nuclear (CBRN) risk scenarios.
🤖 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/devsgnr/bio-safety-peft-lora.laravel-coder-lora-train
Laravel Coder LoRA — end-to-end training data
Instruction-tuning datasets for Laravel Bob (Qwen / CodeLlama / DeepSeek modelfiles).Built from official Laravel docs v10.x–v13.x plus version-detection examples.
Each config has a single schema — do not mix raw (Alpaca) with chat / lora_* (messages).
Configs
Config
Path
Rows
Schema
raw
raw/laravel_training.jsonl
1253
instruction, input, output, topic, version
chat
chat/laravel_training_chat.jsonl
1253… See the full description on the dataset page: https://huggingface.co/datasets/bhavin-gajjar/laravel-coder-lora-train.godot-lora-dataset
Godot LORA Dataset
GDScript training dataset for fine-tuning code models on Godot engine development.
Dataset Info
Total samples: 476
Train split: 428
Validation split: 48
Format: JSONL (instruction, input, output)
Language: GDScript (Godot 4.x)
Languages: German instructions, GDScript code
Sources
godotengine/godot-demo-projects
GDQuest/godot-open-rpg
GDQuest/godot-3d-dodge-the-creeps
bitbrain/beehave (behavior trees)
limboai/limboai (AI for Godot)… See the full description on the dataset page: https://huggingface.co/datasets/matzejo/godot-lora-dataset.ptv3-bericht-lora-de-300
ptv3-bericht-lora-de-300
Synthetic German dataset for fine-tuning LLMs to generate structured psychotherapy reports (PTV-3 / Bericht an den Gutachter) from therapy session transcripts.
Overview
Property
Value
Samples
311 (280 train / 31 val)
Language
German
Format
ChatML JSONL (system / user / assistant)
Teacher model
Qwen2.5-27B (local)
Generation
Two-stage: seed → session transcript → PTV-3 JSON report
Schema
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/speed-brain-ai/ptv3-bericht-lora-de-300.cross_lingual_transfer_dialog_generationcross-lingual transfer in dialog generation
Chinese dialogs in movie domain: Chinese_corpus/train.jsonl, Chinese_corpus/dev.jsonl, Chinese_corpus/test.jsonl. The sizes are 500/50/500.
English dialogs in movie domain: English_corpus/train.jsonl, English_corpus/dev.jsonl. The sizes are 400k/20k.
Chinese dialogs for test in music/book/tech domain: other_domains/music.test.jsonl, other_domains/book.test.jsonl, other_domains/tech.test.jsonl. The sizes are 500/500/500.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/lorashen/cross_lingual_transfer_dialog_generation.JMT-Bench-result_self-rewarding_Mistral-7B-lora
JMT-Bench result
Answer language
JMT-Benchの回答のうち、Englishで回答した件数
Model
Count
mistralai/Mistral-7B-v0.3
25
HachiML/Mistral-7B-v0.3-m1-lora
7
HachiML/Mistral-7B-v0.3-m2-lora
7
HachiML/Mistral-7B-v0.3-m3-lora
2
vedun-lora-data
📚 vedun-lora-data
Синтетический датасет вопрос/ответ для дообучения языковой модели в роли
«древнеславянского ведуна» — она разбирает русские слова через Буквицу
и собирает из значений букв осмысленные ответы.
Использован для обучения серии моделей Qwen3-14B-Vedun-v5 (bf16 / q8 / q4).
⚠️ Дисклеймер. Это развлекательный pet-project, а не пророчества, религия
или философия. Результаты модели — стилизованный текст в эстетике
«древнеславянской буквицы» по мотивам соответствующих… See the full description on the dataset page: https://huggingface.co/datasets/TsitkoD/vedun-lora-data.lora-rules-dataset
LoRA Rules Dataset
Synthetic behavioral rules dataset for training a hypernetwork that generates
LoRA adapters on-the-fly from structured rule strings.
Format
Each record is a JSON line with fields:
rule_id — unique identifier
rule_type — one of: Constraint, Format, Knowledge, Persona, Safety, Tone
weight — float 0.0–1.0, importance of the rule
description — natural language rule description
raw — full rule string [RuleType|Weight] Description
training_examples — list of… See the full description on the dataset page: https://huggingface.co/datasets/broadfield-dev/lora-rules-dataset.lora-rules-qwen3-0.6b-r8-n180
LoRA Rules Dataset
Synthetic behavioral rules dataset for training a hypernetwork that generates
LoRA adapters on-the-fly from structured rule strings.
Format
Each record is a JSON line with fields:
rule_id — unique identifier
rule_type — one of: Constraint, Format, Knowledge, Persona, Safety, Tone
weight — float 0.0–1.0, importance of the rule
description — natural language rule description
raw — full rule string [RuleType|Weight] Description
training_examples — list of… See the full description on the dataset page: https://huggingface.co/datasets/broadfield-dev/lora-rules-qwen3-0.6b-r8-n180.LoRA_Fine-Tune_Q_and_A
Introduction
This dataset is used to fine-tune Deepseek-R1.
