datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
local-agentic-coding-bench-8gb-vram-2026-05
agentic coding benchmark: local LLMs on 8GB VRAM
can local LLMs do agentic coding (multi-turn tool calling, file creation, debugging) on consumer hardware? this dataset captures real test results.
hardware
GPU: NVIDIA RTX 4060 Ti 8GB
CPU: Intel i7-14700F
RAM: 32 GB DDR5
OS: Windows 11 + WSL2 (Ubuntu)
inference: llama-server (turboquant fork of llama.cpp)
what was tested
two agent frameworks:
Hermes Agent (NousResearch): structured tool calling with… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/local-agentic-coding-bench-8gb-vram-2026-05.localagent-dispatch-data
LocalAgent Dispatch Data
Synthetic data for training/evaluating a generable tool-dispatch model over a 50-tool surface
(route head → dense selector → pointer-copy). A static snapshot of the deterministic generators in
LocalAgent (src/localagent/data/). Train/eval are
disjoint in both phrasing and slot values. Companion model + demo:
danelcsb/localagent-tiny-30m-byte ·
Space.
Configs
config
rows (train/eval)
what it is
paraphrase
1000 / 1000
many natural… See the full description on the dataset page: https://huggingface.co/datasets/danelcsb/localagent-dispatch-data.mozgach_localizations
Mozgach Localizations Dataset
Dataset Description
This dataset contains localization strings for the Mozgach application, providing translations from Russian to multiple languages including Chinese, Arabic, and others. The dataset is formatted for instruction-following language models and translation tasks.
Languages
Source Language: Russian (ru)
Target Languages: Chinese (zh), Arabic (ar), and others
Dataset Structure
Each entry in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/nativemind/mozgach_localizations.hybrid_gym_func_localize_raw
Hybrid Gym: Function Localization Dataset
Dataset for the Function Localization benchmark task. Each instance contains a function description (without file path or name), and the agent must locate the function in the repository and add a docstring.
See the benchmark README for usage instructions.
Fields
instance_id: Unique identifier
repo: GitHub repository (owner/repo)
base_commit: Commit hash to checkout
file_path: Path to the file containing the target function… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-gym/hybrid_gym_func_localize_raw.localrule-kr
wellsa-ai localrule-kr
자치법규 (local ordinances, rules).
MiniLex 7-domain Korean lawdata infrastructure.
Snapshot
Documents: 159,910
Snapshot date: 2026-06-05
Source: 법제처 DRF OpenAPI
Pipeline: daily cron 06:15~07:30 KST (scrape → fetch_body → convert → commit)
Schema
Each row in train.jsonl:
field
type
description
id
string
stable document id
name
string
document name (Korean)
category
string
subdirectory (year / type / dept)… See the full description on the dataset page: https://huggingface.co/datasets/wellsa-ai/localrule-kr.birmingham-al-local-businesses
Birmingham, AL Local Business Dataset
Structured data about verified local businesses in the Birmingham, Alabama metropolitan area. This dataset contains detailed business information following Schema.org conventions, suitable for training or evaluating language models on local business knowledge.
Dataset Description
This dataset provides comprehensive structured information about local businesses in Birmingham, Alabama, including:
Business names and alternate names… See the full description on the dataset page: https://huggingface.co/datasets/vltrinkle/birmingham-al-local-businesses.
