Nous Research
hermes-function-calling-v1
Hermes Function-Calling V1
This dataset is the compilation of structured output and function calling data used in the Hermes 2 Pro series of models.
This repository contains a structured output dataset with function-calling conversations, json-mode, agentic json-mode and structured extraction samples, designed to train LLM models in performing function calls and returning structured output based on natural language instructions. The dataset features various conversational scenarios… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/hermes-function-calling-v1.SWE-smith-oracleThis is a version of SWE-bench/SWE-smith filtered for non-empty problem_statement and formatted into the oracle setting of SWE-bench where the files edited by the patch are displayed to the agent. This problem presentation is made available in a text column, following the format of princeton-nlp/SWE-bench_Lite_oracle.
Hermes-3-Dataset
json-mode-evalNousResearch-Hermes-3-Dataset-multiturn
Hermes 3 Multiturn
This is a filtered subset of NousResearch/Hermes-3-Dataset
containing only multiturn conversations with more than three messages.
Conversations with repetitive or trivial replies (for example, repeated "OK") have been excluded to improve quality.
eval-Qwen3-235B-A22B-reasoning
qwen-235b-a22-thinking Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.782
math_pass@1:64_samples
64
0.5%
aime25
0.718
math_pass@1:64_samples
64
0.1%
arenahard
0.939
eval/overall_winrate
500
0.0%
bbh_generative
0.884
extractive_match
1
0.0%
creative-writing-v3
0.775
creative_writing_score
96
0.0%
drop_generative_nous
0.903
drop_acc
1
0.0%
eqbench3
0.800
eqbench_score
135
0.0%
gpqa_diamond
0.697… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Qwen3-235B-A22B-reasoning.
