datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
self-rewarding_AIFT_MSv0.3_lora
self-rewarding_AIFT_MSv0.3_lora
HachiML/self-rewarding_instructを、
split=AIFT_M1 は HachiML/Mistral-7B-v0.3-m1-lora
split=AIFT_M2 は HachiML/Mistral-7B-v0.3-m2-lora
でそれぞれself-rewardingして作成したAIFT(AI Feedback Tuning) dataです。
手順は以下の通りです。
HachiML/self-rewarding_instructのInstructionに対する回答を各モデルで4つずつ作成
回答に対して各モデルで点数評価
最高評価の回答をchosen、最低評価の回答をrejectedとする
詳細はself-rewardingの論文を参照してください。
Dataset Details
Dataset Description
Curated by: HachiMLLanguage(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/HachiML/self-rewarding_AIFT_MSv0.3_lora.JMT-Bench-result_self-rewarding_Mistral-7B-lora
JMT-Bench result
Answer language
JMT-Benchの回答のうち、Englishで回答した件数
Model
Count
mistralai/Mistral-7B-v0.3
25
HachiML/Mistral-7B-v0.3-m1-lora
7
HachiML/Mistral-7B-v0.3-m2-lora
7
HachiML/Mistral-7B-v0.3-m3-lora
2
loracle-ia-diverse-qa-subagent-10q
Loracle IA Diverse QA Subagent 10Q
This dataset is a derived, expanded version of ceselder/loracle-ia-diverse-qa.
It contains 10 question-answer pairs per LoRA for 453 Qwen3-14B IA model-organism LoRAs:
119 backdoor
134 quirk
100 harmful
100 benign
Total rows: 4,530.
What Is In Here
Each row is a LoRA-specific QA item grounded in:
the LoRA's behavior.txt
two selected support prompts from its train.jsonl
a same-family distractor LoRA
a paired mirror LoRA when… See the full description on the dataset page: https://huggingface.co/datasets/japhba/loracle-ia-diverse-qa-subagent-10q.lora-rules-dataset
LoRA Rules Dataset
Synthetic behavioral rules dataset for training a hypernetwork that generates
LoRA adapters on-the-fly from structured rule strings.
Format
Each record is a JSON line with fields:
rule_id — unique identifier
rule_type — one of: Constraint, Format, Knowledge, Persona, Safety, Tone
weight — float 0.0–1.0, importance of the rule
description — natural language rule description
raw — full rule string [RuleType|Weight] Description
training_examples — list of… See the full description on the dataset page: https://huggingface.co/datasets/broadfield-dev/lora-rules-dataset.lora-rules-qwen3-0.6b-r8-n180
LoRA Rules Dataset
Synthetic behavioral rules dataset for training a hypernetwork that generates
LoRA adapters on-the-fly from structured rule strings.
Format
Each record is a JSON line with fields:
rule_id — unique identifier
rule_type — one of: Constraint, Format, Knowledge, Persona, Safety, Tone
weight — float 0.0–1.0, importance of the rule
description — natural language rule description
raw — full rule string [RuleType|Weight] Description
training_examples — list of… See the full description on the dataset page: https://huggingface.co/datasets/broadfield-dev/lora-rules-qwen3-0.6b-r8-n180.
