datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ProjectLucia_Hera
ProjectLucia_Hera
'루시아 발렌타인' 페르소나 LoRA 학습용 한국어 데이터셋. 두 개의 config로 이루어진다.
config
split
행 수
내용
default
train
7,289
단일 턴 한국어 페르소나 대화 (instruction / response)
tools
train / eval
5,495 / 322
도구 호출(function calling) 대화
default
기존과 동일. 루시아 ↔ 멜리사님 1:1 단일 턴 대화. 평균 길이는 질문 46자 → 답변 84자.
학습 시 [system(캐릭터 컨셉), user(instruction), assistant(response)]로 조립해 쓴다.
tools
루시아를 자비스 모드(OpenMascotAI 마스코트가 윈도우를 실제로 조작하는 모드)에서
쓰기 위한 도구 호출 학습 데이터. 페르소나 LoRA를… See the full description on the dataset page: https://huggingface.co/datasets/MelissaJ/ProjectLucia_Hera.Mental-Health-Safety-Eval
Dataset Overview
Created by the HeraFox team, this dataset aims to build awareness for mental health and support research into AI safety and crisis intervention. It evaluates how conversational AI models navigate sensitive self-harm risks, roleplay boundary-blurring, and third-party concerns by delivering safe, empathetic, and resource-connected responses.
Usage & Credits
This dataset is free to use, modify, and distribute for any purpose. While not required, attribution to the HeraFox team… See the full description on the dataset page: https://huggingface.co/datasets/HeraFox-ai/Mental-Health-Safety-Eval.gospel-aloe-hera-v6
Gospel Aloe — Hera duplex training set (v6)
Everyday-conversation companion to
InternalCan/gospel-didactic-hera-v6.
Same codes-only Hera schema (32 Mimi codebooks, word-level timestamps, v6
B-channel corruption). Shards were packed on three nodes and published into
this one repo:
Source
Shard prefix
Role
This 8×H100 box + node A
local-*, nodeA-*
~3.1k conversations
Node B (63.141.33.128)
nodeB-*
~7.3k conversations
sample_id is the join key. There is no overlap… See the full description on the dataset page: https://huggingface.co/datasets/InternalCan/gospel-aloe-hera-v6.Herald_proofsThis is the proof part of the Herald dataset, which consists of 45k NL-FL proofs.
Lean version: leanprover--lean4---v4.11.0
Bibtex citation
@inproceedings{
gao2025herald,
title={Herald: A Natural Language Annotated Lean 4 Dataset},
author={Guoxiong Gao and Yutong Wang and Jiedong Jiang and Qi Gao and Zihan Qin and Tianyi Xu and Bin Dong},
booktitle={The Thirteenth International Conference on Learning Representations},
year={2025},
url={https://openreview.net/forum?id=Se6MgCtRhz}
}
Famous-paintingsContains 100 images of famous paintings you can use for your projects. Ultra-lightweight dataset for quick and simple testing or training
Herald_statementsThis is the statement part of the Herald dataset, which consists of 580k NL-FL statement pairs.
Lean version: leanprover--lean4---v4.11.0
Bibtex citation
@inproceedings{
gao2025herald,
title={Herald: A Natural Language Annotated Lean 4 Dataset},
author={Guoxiong Gao and Yutong Wang and Jiedong Jiang and Qi Gao and Zihan Qin and Tianyi Xu and Bin Dong},
booktitle={The Thirteenth International Conference on Learning Representations},
year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/FrenzyMath/Herald_statements.HERAHERAHellenic Retrieval-Augmented — a long-context RAG benchmark for Greek (retrieval · reader · end-to-end), from Greek Wikipedia
HERA (Hellenic Retrieval-Augmented) is a native-Greek benchmark for long-context retrieval-augmented
generation with citations, abstention, and multi-hop reasoning. Greek is largely absent from
the major multilingual RAG/retrieval benchmarks (MIRACL, Mr.TyDi, mMARCO); this helps fill that gap.
Source: Greek Wikipedia (elwiki latest dump) — CC-BY-SA 4.0
Size: 4,946… See the full description on the dataset page: https://huggingface.co/datasets/KIEFERSA/HERA.us-army-fm-instructThis is a multiturn instruct tuning dataset with 2,333,924 trainable tokens, created with Augmentoolkit, covering the material in the majority of the US Army Field Manuals that are publicly available.
Unlike many previous Augmentoolkit datasets, the questions and answers here are without fluff and are more "to the point". This "sharper" data is intended to help the LLM with recalling facts.
There are three main datasets included here: "vanilla", "negative" and "long".
Vanilla data is simple… See the full description on the dataset page: https://huggingface.co/datasets/Heralax/us-army-fm-instruct.RPToolkit-demo-datasetRPToolkit is a data generation pipeline, part of Augmentoolkit, that generates synthetic RP sessions inspired by input stories. Basically: feed in Lord of the Rings, get out high fantasy adventure RPs.
This dataset, containing over a million trainable tokens across around 1000 RP sessions, is meant to showcase the capabilities of this pipeline.
The input texts used were: a variety of myths and classic stories from Gutenberg; the first few chapters of some miscellaneous webnovels and… See the full description on the dataset page: https://huggingface.co/datasets/Heralax/RPToolkit-demo-dataset.Augmentoolkit-demo
This is a demonstration dataset created using Augmentoolkit and some Project Gutenberg books.
Many of the people who follow me on HF do AI RP, so this dataset was generated with the AI RP mode turned on. Rest assured; there is a professional "Assistant Mode" available for the pipeline
Also the prompt polish on this older version was a bit lacking
Augmentoolkit lets you use local models running on your own machine to create datasets based on any text you can conceive of.… See the full description on the dataset page: https://huggingface.co/datasets/Heralax/Augmentoolkit-demo.test-atk-dataset-do-not-use-3HeraiHench__Phi-4-slerp-ReasoningRP-14B-details
Dataset Card for Evaluation run of HeraiHench/Phi-4-slerp-ReasoningRP-14B
Dataset automatically created during the evaluation run of model HeraiHench/Phi-4-slerp-ReasoningRP-14B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HeraiHench__Phi-4-slerp-ReasoningRP-14B-details.herald-logits
HERALD Logit Signals
Per-token logit-derived signals from Qwen2.5-7B-Instruct generating
under six KV-cache compression methods, plus an uncompressed
baseline. The dataset accompanies the paper HERALD: Hazard
Estimation via Real-time Analysis of Logit Distributions.
The release lets external users reproduce every per-token and
per-run claim in the paper, train alternative predictors against
the same labels, and explore beyond the H = 25 horizon.
Configurations
tokens… See the full description on the dataset page: https://huggingface.co/datasets/Jocana/herald-logits.Augmental-Dataset
A High-Quality AI Augmented Dataset for RP and conversation
This dataset is comprised of lines from the Visual Novel Steins;Gate, which have been filtered, reformatted, AI-rewritten (many of them twice), and in a few cases, manually quality checked.
The flagship model of this dataset (a finetune on top of MythoMax) can be found here!
It contains a large number of RP-focused, multiturn conversational training examples, from the perspectives of multiple characters.
The "Scenario"… See the full description on the dataset page: https://huggingface.co/datasets/Heralax/Augmental-Dataset.he-raw-28BHeraiHench__Marge-Qwen-Math-7B-details
Dataset Card for Evaluation run of HeraiHench/Marge-Qwen-Math-7B
Dataset automatically created during the evaluation run of model HeraiHench/Marge-Qwen-Math-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HeraiHench__Marge-Qwen-Math-7B-details.ProjectLucia_Hera_fullHeraiHench__Double-Down-Qwen-Math-7B-details
Dataset Card for Evaluation run of HeraiHench/Double-Down-Qwen-Math-7B
Dataset automatically created during the evaluation run of model HeraiHench/Double-Down-Qwen-Math-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HeraiHench__Double-Down-Qwen-Math-7B-details.test-atk-dataset-do-not-use
Dataset Card for "test-atk-dataset-do-not-use"
More Information needed
Augmentoolkit-DatagenThis is the dataset used to make the Augmentoolkit data specialist model.
It was made by running Deepseek V3 through all the Augmentoolkit pipelines on a variety of data sources.
antiquated-warfareThis is an instruct tuning dataset with 3 million trainable tokens, created with Augmentoolkit, covering the material in the following Project Gutenberg books:
The Art of War (Sun Tzu)
On War (Clausewitz)
Battle Studies; Ancient and Modern Battle (Charles Jean Jacques Joseph Ardant du Picq)
Elements of Military Art and Science
Blue Shirt and Khaki: A Comparison
Lectures on Land Warfare; A tactical Manual for the Use of Infantry Officers
The Making of a Modern Army and its Operations in the… See the full description on the dataset page: https://huggingface.co/datasets/Heralax/antiquated-warfare.crawl-heraldmalaysiaHeralax-philosophy-instructThis is a multiturn instruct tuning dataset with 729,129 trainable tokens, created with Augmentoolkit, covering the material in the following Project Gutenberg books:
The Problems of Philosophy (Bertrand Russell)
Beyond Good and Evil (Nietzsche)
Thus Spake Zarathustra: A Book for All and None (Nietzsche)
The Prince (Machiavelli)
Second Treatise of Government
These books were chosen simply because they were the top 5 books in the philosophy category on Gutenberg. This is perhaps why at least… See the full description on the dataset page: https://huggingface.co/datasets/MrRobotoAI/Heralax-philosophy-instruct.Heralax-Manners-datasetThis is a multiturn instruct tuning dataset with 1,256,972 trainable tokens, created with Augmentoolkit, covering the material in the following Project Gutenberg books:
Why Etiquette? Because by studying manners, LLMs study human behavior and culture.
Perfect Behavior: A Guide for Ladies and Gentlemen in All Social Crises
The Book of Good Manners; a Guide to Polite Usage for All Social Functions
The Laws of Etiquette; Or, Short Rules and Reflections for Conduct in Society
Manners and Social… See the full description on the dataset page: https://huggingface.co/datasets/MrRobotoAI/Heralax-Manners-dataset.philosophy-instructThis is a multiturn instruct tuning dataset with 729,129 trainable tokens, created with Augmentoolkit, covering the material in the following Project Gutenberg books:
The Problems of Philosophy (Bertrand Russell)
Beyond Good and Evil (Nietzsche)
Thus Spake Zarathustra: A Book for All and None (Nietzsche)
The Prince (Machiavelli)
Second Treatise of Government
These books were chosen simply because they were the top 5 books in the philosophy category on Gutenberg. This is perhaps why at least… See the full description on the dataset page: https://huggingface.co/datasets/Heralax/philosophy-instruct.Mannerstral-datasetThis is a multiturn instruct tuning dataset with 1,256,972 trainable tokens, created with Augmentoolkit, covering the material in the following Project Gutenberg books:
Why Etiquette? Because by studying manners, LLMs study human behavior and culture.
Perfect Behavior: A Guide for Ladies and Gentlemen in All Social Crises
The Book of Good Manners; a Guide to Polite Usage for All Social Functions
The Laws of Etiquette; Or, Short Rules and Reflections for Conduct in Society
Manners and Social… See the full description on the dataset page: https://huggingface.co/datasets/Heralax/Mannerstral-dataset.cmp-facade-samplesherald_proofs_dpo_ready
DPO-Ready Herald Proofs Dataset
This dataset is a modified version of the FrenzyMath/Herald_proofs dataset, specifically restructured for Direct Preference Optimization (DPO) fine-tuning.
herald-proofs-finweb-dpoDPO-Ready Herald Proofs Dataset
This dataset is a modified version of the FrenzyMath/Herald_proofs and lvwerra/stack-exchange-paired dataset, specifically restructured for Direct Preference Optimization (DPO) fine-tuning.
preference-expression-qwen3-8b
Preference Expression Benchmark: Qwen3-8B Base (Condition 1)
Overview
This dataset contains 84,000 responses (8,400 prompts x 10 samples each) from Qwen3-8B base (no RLHF, no DPO, no instruction tuning) to preference-eliciting prompts across 8 domains, sampled at 3 temperatures (0.7, 1.0, 1.5). Each prompt asks the model to express a personal preference, make a subjective judgment, or choose between options.
The dataset also includes 4,410 Claude-scored binary labels… See the full description on the dataset page: https://huggingface.co/datasets/heraldai/preference-expression-qwen3-8b.
