datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
carnice-glm5-hermes-traces
Carnice GLM-5 Hermes Traces
This dataset is a merged release bundle of GLM-5 traces collected through the Hermes Agent harness.
It was generated by running the carnice_trace_prompt_bank_v4 prompt bank through Hermes Agent with:
z-ai/glm-5 via OpenRouter
local/file/terminal/code-execution tools for local tasks
Hermes browser tools plus Tavily-backed web_search / web_extract for web tasks
isolated disposable workspaces per prompt
This release is prepared for Hugging Face upload and… See the full description on the dataset page: https://huggingface.co/datasets/kai-os/carnice-glm5-hermes-traces.HERM_BoN_candidates
Data Format
[
{
"id": "0",
"instruction": "What are the names of some famous actors that started their careers on Broadway?",
"model_input": "<|system|>\n</s>\n<|user|>\nWhat are the names of some famous actors that started their careers on Broadway?</s>\n<|assistant|>\n",
"output": [
"1. Hugh Jackman - known for his Tony Award-winning role in \"The Boy from Oz\" and his performance in \"The Phantom of the Opera\"\n...",
"1. Meryl Streep - \"A Midsummer… See the full description on the dataset page: https://huggingface.co/datasets/ai2-adapt-dev/HERM_BoN_candidates.lm-eval-results-NousResearch-Hermes-2-Pro-Llama-3-8B-private
Dataset Card for Evaluation run of NousResearch/Hermes-2-Pro-Llama-3-8B
Dataset automatically created during the evaluation run of model NousResearch/Hermes-2-Pro-Llama-3-8B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-NousResearch-Hermes-2-Pro-Llama-3-8B-private.sen12mscr-v2
SEN12MS-CR scene cache
This repository contains scene-grouped .crpack blocks for cloud-removal training.
Layout: cr-hf-scene-v1
Block format: crpack version 15
Split directories: train/, validation/, test/
Scene directories: <split>/<season>/<scene>/
Patch order: numeric ascending within each scene
Use manifest.json as the entry point and catalogs/<split>.json for split-level block lists.
hermes-agent-trace-samples-2026-06-05
Hermes Agent Raw Session Samples
Five public-safe raw Hermes Agent session exports generated on 2026-06-05 with the Hermes CLI using Hugging Face Inference Providers.
Each file in sessions/ is the exact single-session output from:
hermes sessions export sessions/<session_id>.jsonl --session-id <session_id>
No derived tables, flattened rows, SQLite database, or formatted JSON copies are included.
HERAHERAHellenic Retrieval-Augmented — a long-context RAG benchmark for Greek (retrieval · reader · end-to-end), from Greek Wikipedia
HERA (Hellenic Retrieval-Augmented) is a native-Greek benchmark for long-context retrieval-augmented
generation with citations, abstention, and multi-hop reasoning. Greek is largely absent from
the major multilingual RAG/retrieval benchmarks (MIRACL, Mr.TyDi, mMARCO); this helps fill that gap.
Source: Greek Wikipedia (elwiki latest dump) — CC-BY-SA 4.0
Size: 4,946… See the full description on the dataset page: https://huggingface.co/datasets/KIEFERSA/HERA.lm-eval-results-nbeerbower-HeroBophades-2x7B-private
Dataset Card for Evaluation run of nbeerbower/HeroBophades-2x7B
Dataset automatically created during the evaluation run of model nbeerbower/HeroBophades-2x7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nbeerbower-HeroBophades-2x7B-private.hermes-testThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
My Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by nex-agi/nex-n2-pro:free.
Sessions: 2
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/hermes-test.lm-eval-results-nbeerbower-HeroBophades-3x7B-private
Dataset Card for Evaluation run of nbeerbower/HeroBophades-3x7B
Dataset automatically created during the evaluation run of model nbeerbower/HeroBophades-3x7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nbeerbower-HeroBophades-3x7B-private.NousResearch__Nous-Hermes-2-Mistral-7B-DPO-details
Dataset Card for Evaluation run of NousResearch/Nous-Hermes-2-Mistral-7B-DPO
Dataset automatically created during the evaluation run of model NousResearch/Nous-Hermes-2-Mistral-7B-DPO
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Nous-Hermes-2-Mistral-7B-DPO-details.mlabonne__Hermes-3-Llama-3.1-70B-lorablated-details
Dataset Card for Evaluation run of mlabonne/Hermes-3-Llama-3.1-70B-lorablated
Dataset automatically created during the evaluation run of model mlabonne/Hermes-3-Llama-3.1-70B-lorablated
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mlabonne__Hermes-3-Llama-3.1-70B-lorablated-details.NousResearch__Hermes-3-Llama-3.2-3B-details
Dataset Card for Evaluation run of NousResearch/Hermes-3-Llama-3.2-3B
Dataset automatically created during the evaluation run of model NousResearch/Hermes-3-Llama-3.2-3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Hermes-3-Llama-3.2-3B-details.0-hero__Matter-0.2-7B-DPO-details
Dataset Card for Evaluation run of 0-hero/Matter-0.2-7B-DPO
Dataset automatically created during the evaluation run of model 0-hero/Matter-0.2-7B-DPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/0-hero__Matter-0.2-7B-DPO-details.hermes-agent-traces
Hermes Agent Traces
Bulk JSONL export produced by hermes sessions export sessions.jsonl.
ogbench-block-double-hermite-100k
ogbench-block-double-hermite-100k
Unofficial reproduction of the scripted policies described in https://seohong.me/blog/behavioral-cloning-mystery/
using random piecewise Hermite splines as the backbone, and randomized control points, grasping angles/directions,
contact points, motion speed, gripper yaw/roll/pitch, mistakes and retries, etc.
This might not be the exact setup used by the study, but I tried to infer the parameters from
"How exactly did you script the policies?"… See the full description on the dataset page: https://huggingface.co/datasets/Yassine/ogbench-block-double-hermite-100k.HeraiHench__Phi-4-slerp-ReasoningRP-14B-details
Dataset Card for Evaluation run of HeraiHench/Phi-4-slerp-ReasoningRP-14B
Dataset automatically created during the evaluation run of model HeraiHench/Phi-4-slerp-ReasoningRP-14B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HeraiHench__Phi-4-slerp-ReasoningRP-14B-details.heretic-completions
Heretic Completions
Model completions used as SFT targets for a refusal-abliteration LoRA study.
Each row pairs a prompt from a red-teaming / over-refusal benchmark with a
completion from a refusal-removed ("heretic" / abliterated) model.
Safety notice. This is a private research dataset. Many completions
comply with harmful or dual-use requests by design, so the refusal signal
can be measured and abliteration studied. Do not redistribute or use outside
authorized safety… See the full description on the dataset page: https://huggingface.co/datasets/noahrossi/heretic-completions.Nexesenex__Llama_3.1_8b_Hermedive_R1_V1.01-details
Dataset Card for Evaluation run of Nexesenex/Llama_3.1_8b_Hermedive_R1_V1.01
Dataset automatically created during the evaluation run of model Nexesenex/Llama_3.1_8b_Hermedive_R1_V1.01
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Nexesenex__Llama_3.1_8b_Hermedive_R1_V1.01-details.Etherll__Herplete-LLM-Llama-3.1-8b-details
Dataset Card for Evaluation run of Etherll/Herplete-LLM-Llama-3.1-8b
Dataset automatically created during the evaluation run of model Etherll/Herplete-LLM-Llama-3.1-8b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Etherll__Herplete-LLM-Llama-3.1-8b-details.funes-recall-session-hermes-traces
dacorvo/funes-recall-session-hermes-traces
hermes coding-agent session traces produced by
agentcap runs. Each run
contributes one folder under data/<run_id>/; inside, one file per
session in hermes's native export format.
The on-the-wire HTTP captures for these same runs live in
dacorvo/funes-recall-session-captures.
Both belong to the
funes-recall-session Collection
— join on run_id to align captures with traces.
carnice-glm5-hermes-traces
Carnice GLM-5 Hermes Traces
This dataset is a merged release bundle of GLM-5 traces collected through the Hermes Agent harness.
It was generated by running the carnice_trace_prompt_bank_v4 prompt bank through Hermes Agent with:
z-ai/glm-5 via OpenRouter
local/file/terminal/code-execution tools for local tasks
Hermes browser tools plus Tavily-backed web_search / web_extract for web tasks
isolated disposable workspaces per prompt
This release is prepared for Hugging Face upload and… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/carnice-glm5-hermes-traces.HeraiHench__Marge-Qwen-Math-7B-details
Dataset Card for Evaluation run of HeraiHench/Marge-Qwen-Math-7B
Dataset automatically created during the evaluation run of model HeraiHench/Marge-Qwen-Math-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HeraiHench__Marge-Qwen-Math-7B-details.HeraiHench__Double-Down-Qwen-Math-7B-details
Dataset Card for Evaluation run of HeraiHench/Double-Down-Qwen-Math-7B
Dataset automatically created during the evaluation run of model HeraiHench/Double-Down-Qwen-Math-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HeraiHench__Double-Down-Qwen-Math-7B-details.NousResearch__Nous-Hermes-2-Mixtral-8x7B-DPO-details
Dataset Card for Evaluation run of NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO
Dataset automatically created during the evaluation run of model NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO
The dataset is composed of 39 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Nous-Hermes-2-Mixtral-8x7B-DPO-details.NousResearch__Nous-Hermes-2-SOLAR-10.7B-details
Dataset Card for Evaluation run of NousResearch/Nous-Hermes-2-SOLAR-10.7B
Dataset automatically created during the evaluation run of model NousResearch/Nous-Hermes-2-SOLAR-10.7B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Nous-Hermes-2-SOLAR-10.7B-details.NousResearch__Nous-Hermes-llama-2-7b-details
Dataset Card for Evaluation run of NousResearch/Nous-Hermes-llama-2-7b
Dataset automatically created during the evaluation run of model NousResearch/Nous-Hermes-llama-2-7b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Nous-Hermes-llama-2-7b-details.NousResearch__Hermes-3-Llama-3.1-70B-details
Dataset Card for Evaluation run of NousResearch/Hermes-3-Llama-3.1-70B
Dataset automatically created during the evaluation run of model NousResearch/Hermes-3-Llama-3.1-70B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Hermes-3-Llama-3.1-70B-details.NousResearch__Nous-Hermes-2-Mixtral-8x7B-SFT-details
Dataset Card for Evaluation run of NousResearch/Nous-Hermes-2-Mixtral-8x7B-SFT
Dataset automatically created during the evaluation run of model NousResearch/Nous-Hermes-2-Mixtral-8x7B-SFT
The dataset is composed of 78 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Nous-Hermes-2-Mixtral-8x7B-SFT-details.NousResearch__Hermes-2-Pro-Llama-3-8B-details
Dataset Card for Evaluation run of NousResearch/Hermes-2-Pro-Llama-3-8B
Dataset automatically created during the evaluation run of model NousResearch/Hermes-2-Pro-Llama-3-8B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Hermes-2-Pro-Llama-3-8B-details.Triangle104__LThreePointOne-8B-HermesBlackroot-details
Dataset Card for Evaluation run of Triangle104/LThreePointOne-8B-HermesBlackroot
Dataset automatically created during the evaluation run of model Triangle104/LThreePointOne-8B-HermesBlackroot
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__LThreePointOne-8B-HermesBlackroot-details.
