CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MelissaJ /ProjectLucia_Hera ProjectLucia_Hera '루시아 발렌타인' 페르소나 LoRA 학습용 한국어 데이터셋. 두 개의 config로 이루어진다. config split 행 수 내용 default train 7,289 단일 턴 한국어 페르소나 대화 (instruction / response) tools train / eval 5,495 / 322 도구 호출(function calling) 대화 default 기존과 동일. 루시아 ↔ 멜리사님 1:1 단일 턴 대화. 평균 길이는 질문 46자 → 답변 84자. 학습 시 [system(캐릭터 컨셉), user(instruction), assistant(response)]로 조립해 쓴다. tools 루시아를 자비스 모드(OpenMascotAI 마스코트가 윈도우를 실제로 조작하는 모드)에서 쓰기 위한 도구 호출 학습 데이터. 페르소나 LoRA를… See the full description on the dataset page: https://huggingface.co/datasets/MelissaJ/ProjectLucia_Hera.texttext-generation10K<n<100K0 likes2.2k downloads22d agoHugging Face02HeraFox-ai /Mental-Health-Safety-Eval Dataset Overview Created by the HeraFox team, this dataset aims to build awareness for mental health and support research into AI safety and crisis intervention. It evaluates how conversational AI models navigate sensitive self-harm risks, roleplay boundary-blurring, and third-party concerns by delivering safe, empathetic, and resource-connected responses. Usage & Credits This dataset is free to use, modify, and distribute for any purpose. While not required, attribution to the HeraFox team… See the full description on the dataset page: https://huggingface.co/datasets/HeraFox-ai/Mental-Health-Safety-Eval.text1K<n<10K10 likes205 downloads27d agoHugging Face03InternalCan /gospel-aloe-hera-v6 Gospel Aloe — Hera duplex training set (v6) Everyday-conversation companion to InternalCan/gospel-didactic-hera-v6. Same codes-only Hera schema (32 Mimi codebooks, word-level timestamps, v6 B-channel corruption). Shards were packed on three nodes and published into this one repo: Source Shard prefix Role This 8×H100 box + node A local-*, nodeA-* ~3.1k conversations Node B (63.141.33.128) nodeB-* ~7.3k conversations sample_id is the join key. There is no overlap… See the full description on the dataset page: https://huggingface.co/datasets/InternalCan/gospel-aloe-hera-v6.tabularautomatic-speech-recognition10K<n<100K1 likes130 downloads9d agoHugging Face04FrenzyMath /Herald_proofsThis is the proof part of the Herald dataset, which consists of 45k NL-FL proofs. Lean version: leanprover--lean4---v4.11.0 Bibtex citation @inproceedings{ gao2025herald, title={Herald: A Natural Language Annotated Lean 4 Dataset}, author={Guoxiong Gao and Yutong Wang and Jiedong Jiang and Qi Gao and Zihan Qin and Tianyi Xu and Bin Dong}, booktitle={The Thirteenth International Conference on Learning Representations}, year={2025}, url={https://openreview.net/forum?id=Se6MgCtRhz} } text10K<n<100K5 likes122 downloads1y agoHugging Face05FrenzyMath /Herald_statementsThis is the statement part of the Herald dataset, which consists of 580k NL-FL statement pairs. Lean version: leanprover--lean4---v4.11.0 Bibtex citation @inproceedings{ gao2025herald, title={Herald: A Natural Language Annotated Lean 4 Dataset}, author={Guoxiong Gao and Yutong Wang and Jiedong Jiang and Qi Gao and Zihan Qin and Tianyi Xu and Bin Dong}, booktitle={The Thirteenth International Conference on Learning Representations}, year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/FrenzyMath/Herald_statements.text100K<n<1M2 likes111 downloads1y agoHugging Face06KIEFERSA /HERAHERAHellenic Retrieval-Augmented — a long-context RAG benchmark for Greek (retrieval · reader · end-to-end), from Greek Wikipedia HERA (Hellenic Retrieval-Augmented) is a native-Greek benchmark for long-context retrieval-augmented generation with citations, abstention, and multi-hop reasoning. Greek is largely absent from the major multilingual RAG/retrieval benchmarks (MIRACL, Mr.TyDi, mMARCO); this helps fill that gap. Source: Greek Wikipedia (elwiki latest dump) — CC-BY-SA 4.0 Size: 4,946… See the full description on the dataset page: https://huggingface.co/datasets/KIEFERSA/HERA.tabularquestion-answering100K<n<1M0 likes102 downloads2mo agoHugging Face07Heralax /us-army-fm-instructThis is a multiturn instruct tuning dataset with 2,333,924 trainable tokens, created with Augmentoolkit, covering the material in the majority of the US Army Field Manuals that are publicly available. Unlike many previous Augmentoolkit datasets, the questions and answers here are without fluff and are more "to the point". This "sharper" data is intended to help the LLM with recalling facts. There are three main datasets included here: "vanilla", "negative" and "long". Vanilla data is simple… See the full description on the dataset page: https://huggingface.co/datasets/Heralax/us-army-fm-instruct.text1K<n<10K12 likes101 downloads2y agoHugging Face08Heralax /RPToolkit-demo-datasetRPToolkit is a data generation pipeline, part of Augmentoolkit, that generates synthetic RP sessions inspired by input stories. Basically: feed in Lord of the Rings, get out high fantasy adventure RPs. This dataset, containing over a million trainable tokens across around 1000 RP sessions, is meant to showcase the capabilities of this pipeline. The input texts used were: a variety of myths and classic stories from Gutenberg; the first few chapters of some miscellaneous webnovels and… See the full description on the dataset page: https://huggingface.co/datasets/Heralax/RPToolkit-demo-dataset.text1K<n<10K17 likes96 downloads2y agoHugging Face09Heralax /test-atk-dataset-do-not-use-3textn<1K0 likes72 downloads2y agoHugging Face10open-llm-leaderboard /HeraiHench__Phi-4-slerp-ReasoningRP-14B-detailsgated Dataset Card for Evaluation run of HeraiHench/Phi-4-slerp-ReasoningRP-14B Dataset automatically created during the evaluation run of model HeraiHench/Phi-4-slerp-ReasoningRP-14B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HeraiHench__Phi-4-slerp-ReasoningRP-14B-details.tabular10K<n<100K0 likes49 downloads2y agoHugging Face11Jocana /herald-logits HERALD Logit Signals Per-token logit-derived signals from Qwen2.5-7B-Instruct generating under six KV-cache compression methods, plus an uncompressed baseline. The dataset accompanies the paper HERALD: Hazard Estimation via Real-time Analysis of Logit Distributions. The release lets external users reproduce every per-token and per-run claim in the paper, train alternative predictors against the same labels, and explore beyond the H = 25 horizon. Configurations tokens… See the full description on the dataset page: https://huggingface.co/datasets/Jocana/herald-logits.tabularother1M<n<10M0 likes44 downloads4mo agoHugging Face12Heralax /Augmental-Dataset A High-Quality AI Augmented Dataset for RP and conversation This dataset is comprised of lines from the Visual Novel Steins;Gate, which have been filtered, reformatted, AI-rewritten (many of them twice), and in a few cases, manually quality checked. The flagship model of this dataset (a finetune on top of MythoMax) can be found here! It contains a large number of RP-focused, multiturn conversational training examples, from the perspectives of multiple characters. The "Scenario"… See the full description on the dataset page: https://huggingface.co/datasets/Heralax/Augmental-Dataset.text1K<n<10K26 likes41 downloads3y agoHugging Face13Beetle-Data /he-raw-28Btabular10M<n<100M0 likes36 downloads4mo agoHugging Face14open-llm-leaderboard /HeraiHench__Marge-Qwen-Math-7B-detailsgated Dataset Card for Evaluation run of HeraiHench/Marge-Qwen-Math-7B Dataset automatically created during the evaluation run of model HeraiHench/Marge-Qwen-Math-7B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HeraiHench__Marge-Qwen-Math-7B-details.tabular10K<n<100K0 likes35 downloads2y agoHugging Face15MelissaJ /ProjectLucia_Hera_fulltext1K<n<10K0 likes34 downloads3mo agoHugging Face16open-llm-leaderboard /HeraiHench__Double-Down-Qwen-Math-7B-detailsgated Dataset Card for Evaluation run of HeraiHench/Double-Down-Qwen-Math-7B Dataset automatically created during the evaluation run of model HeraiHench/Double-Down-Qwen-Math-7B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HeraiHench__Double-Down-Qwen-Math-7B-details.tabular10K<n<100K0 likes32 downloads2y agoHugging Face17Heralax /test-atk-dataset-do-not-use Dataset Card for "test-atk-dataset-do-not-use" More Information needed textn<1K0 likes31 downloads2y agoHugging Face18Heralax /antiquated-warfareThis is an instruct tuning dataset with 3 million trainable tokens, created with Augmentoolkit, covering the material in the following Project Gutenberg books: The Art of War (Sun Tzu) On War (Clausewitz) Battle Studies; Ancient and Modern Battle (Charles Jean Jacques Joseph Ardant du Picq) Elements of Military Art and Science Blue Shirt and Khaki: A Comparison Lectures on Land Warfare; A tactical Manual for the Use of Infantry Officers The Making of a Modern Army and its Operations in the… See the full description on the dataset page: https://huggingface.co/datasets/Heralax/antiquated-warfare.text10K<n<100K7 likes25 downloads2y agoHugging Face19MrRobotoAI /Heralax-philosophy-instructThis is a multiturn instruct tuning dataset with 729,129 trainable tokens, created with Augmentoolkit, covering the material in the following Project Gutenberg books: The Problems of Philosophy (Bertrand Russell) Beyond Good and Evil (Nietzsche) Thus Spake Zarathustra: A Book for All and None (Nietzsche) The Prince (Machiavelli) Second Treatise of Government These books were chosen simply because they were the top 5 books in the philosophy category on Gutenberg. This is perhaps why at least… See the full description on the dataset page: https://huggingface.co/datasets/MrRobotoAI/Heralax-philosophy-instruct.text1K<n<10K0 likes18 downloads2y agoHugging Face20MrRobotoAI /Heralax-Manners-datasetThis is a multiturn instruct tuning dataset with 1,256,972 trainable tokens, created with Augmentoolkit, covering the material in the following Project Gutenberg books: Why Etiquette? Because by studying manners, LLMs study human behavior and culture. Perfect Behavior: A Guide for Ladies and Gentlemen in All Social Crises The Book of Good Manners; a Guide to Polite Usage for All Social Functions The Laws of Etiquette; Or, Short Rules and Reflections for Conduct in Society Manners and Social… See the full description on the dataset page: https://huggingface.co/datasets/MrRobotoAI/Heralax-Manners-dataset.text1K<n<10K0 likes18 downloads2y agoHugging Face21Heralax /Mannerstral-datasetThis is a multiturn instruct tuning dataset with 1,256,972 trainable tokens, created with Augmentoolkit, covering the material in the following Project Gutenberg books: Why Etiquette? Because by studying manners, LLMs study human behavior and culture. Perfect Behavior: A Guide for Ladies and Gentlemen in All Social Crises The Book of Good Manners; a Guide to Polite Usage for All Social Functions The Laws of Etiquette; Or, Short Rules and Reflections for Conduct in Society Manners and Social… See the full description on the dataset page: https://huggingface.co/datasets/Heralax/Mannerstral-dataset.text1K<n<10K2 likes16 downloads2y agoHugging Face22korazer /herald_proofs_dpo_ready DPO-Ready Herald Proofs Dataset This dataset is a modified version of the FrenzyMath/Herald_proofs dataset, specifically restructured for Direct Preference Optimization (DPO) fine-tuning. text10K<n<100K0 likes12 downloads1y agoHugging Face23korazer /herald-proofs-finweb-dpoDPO-Ready Herald Proofs Dataset This dataset is a modified version of the FrenzyMath/Herald_proofs and lvwerra/stack-exchange-paired dataset, specifically restructured for Direct Preference Optimization (DPO) fine-tuning. text10K<n<100K0 likes10 downloads1y agoHugging Face24electricsheepafrica /africa-herams-cabo-delgado-dataset Mozambique: Health Facilities | Africa (original) Size category: n<1K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Health datasets help researchers examine disease burden, service… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-herams-cabo-delgado-dataset.tabulartabular-classificationn<1K0 likes10 downloads1mo agoHugging Face25herambpatil2004 /EANDTCtextn<1K0 likes6 downloads2y agoHugging Face26heranock /logotbltext10K<n<100K0 likes5 downloads2y agoHugging Face27PJMixers-Dev /Heralax_RPToolkit-demo-datasettextn<1K0 likes4 downloads2y agoHugging Face28IkuyoIkuyo /Herald_with_proofgatedtext10K<n<100K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.