nousresearch
hermes-function-calling-v1
Hermes Function-Calling V1
This dataset is the compilation of structured output and function calling data used in the Hermes 2 Pro series of models.
This repository contains a structured output dataset with function-calling conversations, json-mode, agentic json-mode and structured extraction samples, designed to train LLM models in performing function calls and returning structured output based on natural language instructions. The dataset features various conversational scenarios… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/hermes-function-calling-v1.SWE-smith-oracleThis is a version of SWE-bench/SWE-smith filtered for non-empty problem_statement and formatted into the oracle setting of SWE-bench where the files edited by the patch are displayed to the agent. This problem presentation is made available in a text column, following the format of princeton-nlp/SWE-bench_Lite_oracle.
json-mode-evalHermes-3-Dataset
CharacterCodex
Dataset Card for Character Codex
Dataset Summary
The Character Codex is a comprehensive dataset featuring popular characters from a wide array of media types and genres. Each entry includes detailed information about the character, the media source, and a unique scenario involving the character. This dataset is valuable for synthetic data, RAG for generative AI, writers, game developers, and fans who want to explore and utilize rich character descriptions for various… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/CharacterCodex.eval-Qwen3-235B-A22B-reasoning
qwen-235b-a22-thinking Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.782
math_pass@1:64_samples
64
0.5%
aime25
0.718
math_pass@1:64_samples
64
0.1%
arenahard
0.939
eval/overall_winrate
500
0.0%
bbh_generative
0.884
extractive_match
1
0.0%
creative-writing-v3
0.775
creative_writing_score
96
0.0%
drop_generative_nous
0.903
drop_acc
1
0.0%
eqbench3
0.800
eqbench_score
135
0.0%
gpqa_diamond
0.697… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Qwen3-235B-A22B-reasoning.
