datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HERBench
HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering
A challenging benchmark for evaluating multi-evidence integration capabilities of vision-language models
🎉 HERBench has been accepted to CVPR 2026!
🆕 New: Lite-v2 config. We released a refreshed lite_v2 version of the
Lite split (1,971 questions / 68 videos) in which 9 of the 12 tasks were
regenerated and went through additional manual refinement for higher
quality, while TSO, SVA… See the full description on the dataset page: https://huggingface.co/datasets/DanBenAmi/HERBench.zo-bible
zo-bible
A sentence-aligned parallel Bible corpus covering 8 closely related Zo speech varieties and English across 10 translation versions (30,974 canonical verses).
Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev).
Overview
Languages: Hmar (hmr), Mizo (lus), Paite (pck), Vaiphei (vap), Thadou (tcz), Gangte (gnb), Zou (zom), English (eng)
Family: Zo Languages
Volume: 30,974 verse anchors across 66 canonical books (10 translation editions)… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/zo-bible.deepseek-hermes-reasoning-traces
DeepSeek V4 Pro Hermes Reasoning Traces
19,331 multi-turn ChatML + Hermes reasoning traces generated by DeepSeek V4 Pro. Designed for LoRA fine-tuning local models to operate as Hermes Agent instances.
Quick Start
\
Splits
Split
Traces
train
16,431
valid
1,933
test
967
Variants (VRAM-Tiered)
Variant
Max Tokens
Traces
GPU
nano
2,048
15,948
Dev / 7B
budget
4,096
2,149
48GB
standard
8,192
990
64GB
spark
16,384
244… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/deepseek-hermes-reasoning-traces.Herculean
Herculean: A Financial Agentic Benchmark
An offline evaluation benchmark for LLM agents performing five financial-analysis
tasks: trading, hedging, report generation, report evaluation, and
XBRL filing auditing. All tasks run against fully offline data — no live
market or web access — so runs are reproducible and bit-comparable across
models.
This release ships two artifacts:
env.duckdb — a single DuckDB database (and equivalent Parquet files
under data/) containing daily prices… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/Herculean.eval-Hermes-4-14B-reasoning
h4-14b-more-stage1-reasoning Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.554
math_pass@1:64_samples
64
0.1%
aime25
0.468
math_pass@1:64_samples
64
0.1%
arenahard
0.830
eval/overall_winrate
500
0.0%
bbh_generative
0.844
extractive_match
1
0.0%
creative-writing-v3
0.616
creative_writing_score
96
0.0%
drop_generative_nous
0.845
drop_acc
1
0.0%
eqbench3
0.772
eqbench_score
135
0.0%
gpqa_diamond
0.602… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-14B-reasoning.herg_central_inhib
Dataset Details
Dataset Description
Human ether-à-go-go related gene (hERG) is crucial for the coordination
of the heart's beating. Thus, if a drug blocks the hERG, it could lead to severe
adverse effects. Therefore, reliable prediction of hERG liability in the early
stages of drug design is quite important to reduce the risk of cardiotoxicity-related
attritions in the later development stages. There are three targets: hERG_at_1microM,
hERG_at_10microM, and herg_inhib.… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/herg_central_inhib.hero_run_4_math_codeeval-Hermes-4-405B-reasoning
405b-e3-40k-reasoning Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.819
math_pass@1:64_samples
64
5.6%
aime25
0.781
math_pass@1:64_samples
64
5.3%
arenahard
0.937
eval/overall_winrate
500
0.0%
bbh_generative
0.863
extractive_match
1
4.7%
creative-writing-v3
0.793
creative_writing_score
96
0.0%
drop_generative_nous
0.835
drop_acc
1
1.6%
eqbench3
0.855
eqbench_score
135
0.0%
gpqa_diamond
0.706… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-405B-reasoning.eval-Hermes-4-70B-nonreasoning
hermes-70b-nonreasoning Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.095
math_pass@1:64_samples
64
99.4%
aime25
0.073
math_pass@1:64_samples
64
98.2%
arenahard
0.568
eval/overall_winrate
500
0.0%
bbh_generative
0.805
extractive_match
1
100.0%
creative-writing-v3
0.491
creative_writing_score
96
0.0%
drop_generative_nous
0.784
drop_acc
1
100.0%
eqbench3
0.739
eqbench_score
135
0.0%
gpqa_diamond
0.333… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-70B-nonreasoning.herg_karim_et_al
Dataset Details
Dataset Description
A integrated Ether-à-go-go-related gene (hERG) dataset consisting
of molecular structures labelled as hERG (<10uM) and non-hERG (>=10uM) blockers in
the form of SMILES strings was obtained from the DeepHIT, the BindingDB database,
ChEMBL bioactivity database, and other literature.
Curated by:
License: CC BY 4.0
Dataset Sources
corresponding publication
Data source
Citation
BibTeX:
@article{Karim2021,
doi =… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/herg_karim_et_al.eval-Hermes-4-405B-nonreasoning
h4-405b-e3-nonthinking Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.114
math_pass@1:64_samples
64
100.0%
aime25
0.106
math_pass@1:64_samples
64
100.0%
arenahard
0.535
eval/overall_winrate
500
0.0%
bbh_generative
0.687
extractive_match
1
100.0%
creative-writing-v3
0.506
creative_writing_score
96
0.0%
drop_generative_nous
0.776
drop_acc
1
100.0%
eqbench3
0.746
eqbench_score
135
0.0%
gpqa_diamond
0.394… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-405B-nonreasoning.eval-Hermes-4-14B-reasoning-old
h4-e3-overlong-masked-30k-rerun Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.527
math_pass@1:64_samples
64
6.6%
aime25
0.414
math_pass@1:64_samples
64
8.1%
arenahard
0.782
eval/overall_winrate
500
0.0%
bbh_generative
0.844
extractive_match
1
5.8%
creative-writing-v3
0.617
creative_writing_score
96
0.0%
drop_generative_nous
0.827
drop_acc
1
2.4%
eqbench3
0.805
eqbench_score
135
0.0%
gpqa_diamond
0.556… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-14B-reasoning-old.eval-Hermes-4.3-36B
36bpsychev2 Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.719
math_pass@1:64_samples
64
17.6%
aime25
0.693
math_pass@1:64_samples
64
18.8%
bbh_generative
0.864
extractive_match
1
4.8%
drop_generative_nous
0.835
drop_acc
1
2.7%
gpqa_diamond
0.655
gpqa_pass@1:8_samples
8
2.2%
ifeval
0.779
inst_level_loose_acc
1
7.8%
math_500
0.938
math_pass@1:4_samples
4
1.8%
mmlu_generative
0.877
extractive_match1
0.1%
mmlu_pro… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4.3-36B.eval-Hermes-4-14B-nonreasoning-old
h4-14b-nonreasoning-30k-cot Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.105
math_pass@1:64_samples
64
99.7%
aime25
0.066
math_pass@1:64_samples
64
100.0%
arenahard
0.498
eval/overall_winrate
500
0.0%
bbh_generative
0.632
extractive_match
1
100.0%
creative-writing-v3
0.405
creative_writing_score
96
0.0%
drop_generative_nous
0.714
drop_acc
1
100.0%
eqbench3
0.690
eqbench_score
135
0.0%
gpqa_diamond
0.450… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-14B-nonreasoning-old.eval-Hermes-4-14B-nonreasoning
h4-14b-more-stage1-nonreasoning Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.110
math_pass@1:64_samples
64
99.9%
aime25
0.069
math_pass@1:64_samples
64
99.6%
arenahard
0.502
eval/overall_winrate
500
0.0%
bbh_generative
0.740
extractive_match
1
100.0%
creative-writing-v3
0.355
creative_writing_score
96
0.0%
drop_generative_nous
0.739
drop_acc
1
100.0%
eqbench3
0.580
eqbench_score
135
0.0%
gpqa_diamond
0.390… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-14B-nonreasoning.place_bottleeval-Hermes-4-70B-reasoning
hermes-4-70b-reasoning-40k Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.735
math_pass@1:64_samples
64
8.4%
aime25
0.674
math_pass@1:64_samples
64
9.6%
arenahard
0.901
eval/overall_winrate
500
0.0%
bbh_generative
0.878
extractive_match
1
4.8%
creative-writing-v3
0.775
creative_writing_score
96
0.0%
drop_generative_nous
0.850
drop_acc
1
1.4%
eqbench3
0.847
eqbench_score
135
0.0%
gpqa_diamond
0.661… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-70B-reasoning.gallica_heritagetask2_fixpos_200
task2_fixpos_200
Teleoperated Isaac Sim demonstrations for the EBiM benchmark, Task 2,
recorded in the LeRobot v3 dataset format.
Teleoperated using keyboard.
Robot positions are overriden:
robot_x: 2.1
robot_y: 3.05
robot_z: 0.0
robot_yaw: -90.0
Tasks
Pick up the thermal pad and place it on the target RAM board.
At a glance
Codebase version
v3.0
Robot type
fr3duo_mobile_task2
Episodes
200
Frames
174719
FPS
30
Action dim… See the full description on the dataset page: https://huggingface.co/datasets/hermanprawiro/task2_fixpos_200.open-hermes-2.5-sft-mixture-llama3-inference-retrieval-tokensHerbarium-2022-FGVC9_masked
Herbarium 2022 FGVC9 Masked
Segmentation masks for the Herbarium 2022 FGVC9 dataset, stored as RLE-encoded masks in a single Parquet file.
Note: This file covers 15,992 images (63 of 400 shards processed so far).
File
File
Description
masks.parquet
15,992 rows — one per image — with RLE mask, score, species label, and file_name
Schema
Column
Type
Description
dataset
str
Always Herbarium-2022-FGVC9
text_prompt
str
Text… See the full description on the dataset page: https://huggingface.co/datasets/kaityc06/Herbarium-2022-FGVC9_masked.herg_blockers
Dataset Details
Dataset Description
Human ether-à-go-go related gene (hERG) is crucial for the coordination
of the heart's beating. Thus, if a drug blocks the hERG, it could lead to severe
adverse effects. Therefore, reliable prediction of hERG liability in the early
stages of drug design is quite important to reduce the risk of cardiotoxicity
related attritions in the later development stages.
Curated by:
License: CC BY 4.0
Dataset Sources
corresponding… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/herg_blockers.hero_run_4_codeopenthoughts3_herorun_ckpt06500_eval_27e9
mlfoundations-dev/openthoughts3_herorun_ckpt06500_eval_27e9
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 59.95% ± 0.79%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
59.30%
303
511
2
62.82%
321
511
3
60.08%
307
511
4
58.51%
299
511
5
61.45%
314
511
6
57.53%
294
511
eval-Hermes-4.3-36B-centralized
36btorchtitan Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.706
math_pass@1:64_samples
64
24.6%
aime25
0.669
math_pass@1:64_samples
64
26.8%
bbh_generative
0.847
extractive_match
1
9.3%
drop_generative_nous
0.817
drop_acc
1
6.9%
gpqa_diamond
0.649
gpqa_pass@1:8_samples
8
4.2%
ifeval
0.740
inst_level_loose_acc
1
13.1%
math_500
0.923
math_pass@1:4_samples
4
3.1%
mmlu_generative
0.866
extractive_match1
1.4%… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4.3-36B-centralized.carnice-glm5-hermes-traces
Carnice GLM-5 Hermes Traces
This dataset is a merged release bundle of GLM-5 traces collected through the Hermes Agent harness.
It was generated by running the carnice_trace_prompt_bank_v4 prompt bank through Hermes Agent with:
z-ai/glm-5 via OpenRouter
local/file/terminal/code-execution tools for local tasks
Hermes browser tools plus Tavily-backed web_search / web_extract for web tasks
isolated disposable workspaces per prompt
This release is prepared for Hugging Face upload and… See the full description on the dataset page: https://huggingface.co/datasets/kai-os/carnice-glm5-hermes-traces.HERM_BoN_candidates
Data Format
[
{
"id": "0",
"instruction": "What are the names of some famous actors that started their careers on Broadway?",
"model_input": "<|system|>\n</s>\n<|user|>\nWhat are the names of some famous actors that started their careers on Broadway?</s>\n<|assistant|>\n",
"output": [
"1. Hugh Jackman - known for his Tony Award-winning role in \"The Boy from Oz\" and his performance in \"The Phantom of the Opera\"\n...",
"1. Meryl Streep - \"A Midsummer… See the full description on the dataset page: https://huggingface.co/datasets/ai2-adapt-dev/HERM_BoN_candidates.Simia-Tau-SFT-90k-Hermes
🐒 Simia-Tau-SFT-90k-Hermes:
Simia-Tau-SFT-90k-Hermes is the fully synthetic tool-agent dataset, designed to advance tool use for τ² Airline and Retail. It comprises over 90k trajectories synthesized from 5k real-world seed trajectories. The tool use format is Hermes. Models fine-tuned on Simia-Tau-SFT-90k-Hermes outperform much larger closed-source counterparts on the τ² Airline and Retail benchmark.
📄 Technical Report - Discover the methodology and technical details behind this… See the full description on the dataset page: https://huggingface.co/datasets/Simia-Agent/Simia-Tau-SFT-90k-Hermes.lm-eval-results-NousResearch-Hermes-2-Pro-Llama-3-8B-private
Dataset Card for Evaluation run of NousResearch/Hermes-2-Pro-Llama-3-8B
Dataset automatically created during the evaluation run of model NousResearch/Hermes-2-Pro-Llama-3-8B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-NousResearch-Hermes-2-Pro-Llama-3-8B-private.sen12mscr-v2
SEN12MS-CR scene cache
This repository contains scene-grouped .crpack blocks for cloud-removal training.
Layout: cr-hf-scene-v1
Block format: crpack version 15
Split directories: train/, validation/, test/
Scene directories: <split>/<season>/<scene>/
Patch order: numeric ascending within each scene
Use manifest.json as the entry point and catalogs/<split>.json for split-level block lists.
