datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Hermes-3-Dataset
hermes3-uk
Dataset Card for Hermes 3 Ukrainian Fixed Conversations
Dataset Description
Dataset Summary
hermes3-uk-fixed is a Ukrainian translation of the [NousResearch/Hermes-3-Dataset]. The translation was produced with Gemma 3 27B (instruction-tuned). During preparation we removed all system prompts and normalized the message roles and content to match the common schema we use across our dialog datasets.
Languages
Ukrainian (uk)
Dataset Structure
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/hermes3-uk.hermes3-en-fixed
Dataset Card for Hermes 3 Fixed Conversations
Dataset Description
Dataset Summary
hermes3-en-fixed is a [NousResearch/Hermes-3-Dataset]. During preparation we removed all system prompts and normalized the message roles and content to match the common schema we use across our dialog datasets.
Languages
English (en)
Dataset Structure
Data Fields
conversations: list of messages in a dialog (array of objects)
from: normalized sender role — user or assistant… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/hermes3-en-fixed.Hermes-3-Dataset-enPurified-openai-messages
Dataset Card: enPurified
This dataset was updated on January 17th, 2026 to convert the messages from sharegpt to openai messages format. I forgot to include that in the January 13th re-do.
This dataset was updated on January 13th, 2026 to strip out even more math/code. The pruning process reduced the dataset from 958,829 to 117,877 rows of high-quality English prose.
(The script used for this process is uploaded in the files section)
Purpose
The enPurified collection is… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/Hermes-3-Dataset-enPurified-openai-messages.hermes-3-openai-formatfdm-40ch-fresh-hermes3hermes-3-dataset-ru-translated-prompts
Переведенные промты из hermes-3-dataset
Модель-переводчик Gemma-3-27b-it.
Переведены все промты.
Multi-turn промты переведены с учетом контекста англоязычного ответа.
Будет полезно для создания крупных русскоязычных инструктивных датасетов или Online RL.
Translated prompts from hermes-3-dataset
Translator model: Gemma-3-27b-it.
All prompts have been translated.
Multi-turn prompts were translated considering the context of the English response.
This will be useful… See the full description on the dataset page: https://huggingface.co/datasets/kristaller486/hermes-3-dataset-ru-translated-prompts.hermes3_system_promptshermes3_special_system_prompts
Hermes 3 SFT Special System Prompts
Version of the NousResearch/Hermes-3-Dataset dataset used in AMALIA's Supervised Fine-Tuning stage.
This subset was manually selected in order to retain only entries containing custom system prompts that substantially modify the model’s behavior. It were also built splits with the highest quality entries and a translated part of those entries to European Portuguese. Both the quality classification and translation were done using… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/hermes3_special_system_prompts.Hermes-3-Dataset
hermes-3
Converted to be directly supported in MLX-LM and MLX-LM-LoRA.
Example with MLX-LM-LoRA:
mlx_lm_lora.train \
--model mlx-community/Josiefied-Qwen2.5-0.5B-Instruct-abliterated-v1 \
--train \
--data mlx-community/hermes-3-mlx \
--iters 100 \
--max-seq-length 4096
Example with MLX-LM:
mlx_lm.lora \
--model mlx-community/Josiefied-Qwen2.5-0.5B-Instruct-abliterated-v1 \
--train \
--data mlx-community/hermes-3-mlx \
--iters 100 \
--max-seq-length 4096
Hermes-3-Dataset-For-Josieficationwill be used to rewrite the responses into Josie style and create a preference dataset
Hermes-3-shuffledHermes-3hermes3-curated-20kHermes-3-filtered-test2Hermes-3-no-sys-msgHermes-3-Dataset-MLXhermes3_chatmlTriangle104__Hermes3-L3.1-DirtyHarry-8B-details
Dataset Card for Evaluation run of Triangle104/Hermes3-L3.1-DirtyHarry-8B
Dataset automatically created during the evaluation run of model Triangle104/Hermes3-L3.1-DirtyHarry-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Hermes3-L3.1-DirtyHarry-8B-details.Hermes-3-Dataset-ConvertedHermes-3-short-conversationshermes3-quick-probes-multilingual
Hermes3 Quick Probes (Multilingual, Reasoning ON/OFF)
Міні-датасет (20 прикладів) для швидкої перевірки Hermes-3 у режимах reasoning ON/OFF (UA/ES/EN/ID).
Ціль — легкі sanity-checks: де потрібне міркування, а де достатньо стислої відповіді.
Формат
Файл: data.jsonl, по 1 JSON-об’єкту на рядок з полями:
id (string) — унікальний ідентифікатор
lang (uk|es|en|id)
reasoning ("on"|"off")
prompt (string)
expect (dict, опційно: keywords/max_sentences/answer)
Як… See the full description on the dataset page: https://huggingface.co/datasets/segs/hermes3-quick-probes-multilingual.Hermes-3-Dataset
Hermes-3-Llama-3.1-8B-details
Dataset Card for Evaluation run of NousResearch/Hermes-3-Llama-3.1-8B
Dataset automatically created during the evaluation run of model NousResearch/Hermes-3-Llama-3.1-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Hermes-3-Llama-3.1-8B-details.
