hera
Datasets
All datasets matching “hera”ProjectLucia_Hera
ProjectLucia_Hera
'루시아 발렌타인' 페르소나 LoRA 학습용 한국어 데이터셋. 두 개의 config로 이루어진다.
config
split
행 수
내용
default
train
7,289
단일 턴 한국어 페르소나 대화 (instruction / response)
tools
train / eval
5,495 / 322
도구 호출(function calling) 대화
default
기존과 동일. 루시아 ↔ 멜리사님 1:1 단일 턴 대화. 평균 길이는 질문 46자 → 답변 84자.
학습 시 [system(캐릭터 컨셉), user(instruction), assistant(response)]로 조립해 쓴다.
tools
루시아를 자비스 모드(OpenMascotAI 마스코트가 윈도우를 실제로 조작하는 모드)에서
쓰기 위한 도구 호출 학습 데이터. 페르소나 LoRA를… See the full description on the dataset page: https://huggingface.co/datasets/MelissaJ/ProjectLucia_Hera.Mental-Health-Safety-Eval
Dataset Overview
Created by the HeraFox team, this dataset aims to build awareness for mental health and support research into AI safety and crisis intervention. It evaluates how conversational AI models navigate sensitive self-harm risks, roleplay boundary-blurring, and third-party concerns by delivering safe, empathetic, and resource-connected responses.
Usage & Credits
This dataset is free to use, modify, and distribute for any purpose. While not required, attribution to the HeraFox team… See the full description on the dataset page: https://huggingface.co/datasets/HeraFox-ai/Mental-Health-Safety-Eval.gospel-aloe-hera-v6
Gospel Aloe — Hera duplex training set (v6)
Everyday-conversation companion to
InternalCan/gospel-didactic-hera-v6.
Same codes-only Hera schema (32 Mimi codebooks, word-level timestamps, v6
B-channel corruption). Shards were packed on three nodes and published into
this one repo:
Source
Shard prefix
Role
This 8×H100 box + node A
local-*, nodeA-*
~3.1k conversations
Node B (63.141.33.128)
nodeB-*
~7.3k conversations
sample_id is the join key. There is no overlap… See the full description on the dataset page: https://huggingface.co/datasets/InternalCan/gospel-aloe-hera-v6.Herald_proofsThis is the proof part of the Herald dataset, which consists of 45k NL-FL proofs.
Lean version: leanprover--lean4---v4.11.0
Bibtex citation
@inproceedings{
gao2025herald,
title={Herald: A Natural Language Annotated Lean 4 Dataset},
author={Guoxiong Gao and Yutong Wang and Jiedong Jiang and Qi Gao and Zihan Qin and Tianyi Xu and Bin Dong},
booktitle={The Thirteenth International Conference on Learning Representations},
year={2025},
url={https://openreview.net/forum?id=Se6MgCtRhz}
}
Famous-paintingsContains 100 images of famous paintings you can use for your projects. Ultra-lightweight dataset for quick and simple testing or training
Herald_statementsThis is the statement part of the Herald dataset, which consists of 580k NL-FL statement pairs.
Lean version: leanprover--lean4---v4.11.0
Bibtex citation
@inproceedings{
gao2025herald,
title={Herald: A Natural Language Annotated Lean 4 Dataset},
author={Guoxiong Gao and Yutong Wang and Jiedong Jiang and Qi Gao and Zihan Qin and Tianyi Xu and Bin Dong},
booktitle={The Thirteenth International Conference on Learning Representations},
year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/FrenzyMath/Herald_statements.
