datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hermes-pairing-bench
Hermes Pairing — Agentic Benchmark for Local LLMs (Phase A + B)
How well does a local LLM drive an agent? This dataset holds results for pairing local models with
Hermes Agent (NousResearch) — a CodeAct agent: the model
acts by writing Python (execute_code) that orchestrates tools, not by emitting JSON function calls.
Generated with llm-bench-rig on an NVIDIA RTX 5090 (32GB),
llama.cpp / GGUF, under Hermes's real ~3.5K-token system prompt.
Phase A (synthetic). A reproducible… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/hermes-pairing-bench.herman-json-mode
Herman: Indonesian Single-Turn JSON Mode
Herman is an Indonesian language dataset specifically designed
for training LLMs using a single-turn JSON mode. This dataset
is used in Supervised Fine-Tuning (SFT) to improve JSON parsing
capabilities in LLMs. Herman was obtained from Hermes and translated
into Indonesian for the purpose of training Indonesian language models.
Code used for constructing Herman can be found here.
Schema Format
The desired JSON schema can… See the full description on the dataset page: https://huggingface.co/datasets/SulthanAbiyyu/herman-json-mode.gbv-anno
