datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sampled-local-resumes
sampled-local-resumes
This dataset contains synthetic resume data sampled from local folders (20% sample from each folder).
License
This dataset is released under the Apache License 2.0. Please see the LICENSE and NOTICE files for details.
Attribution
Copyright 2025 Fairly AI Inc. dba Asenion
This dataset includes data released by Fairly AI Inc. dba Asenion under the Apache License, Version 2.0.
You may obtain a copy of the License at:… See the full description on the dataset page: https://huggingface.co/datasets/asenion-ai/sampled-local-resumes.mytown-local-gov-meetings
MyTown — open dataset of US & Canadian local-government meetings
The documents themselves, not just the metadata. Most civic datasets publish meeting
titles, dates and links. This one publishes 2,109,683 full text extractions of
the primary documents — the actual agendas and minutes, pulled out of the PDFs — alongside
11,949,495 per-member roll-call votes and 61,661,080 campaign-finance
transactions, all joinable on the same keys.
That combination is the point: you can go from… See the full description on the dataset page: https://huggingface.co/datasets/jazzypajamas/mytown-local-gov-meetings.lca-bug-localization
🏟️ Long Code Arena (Bug localization)
This is the benchmark for the Bug localization task as part of the
🏟️ Long Code Arena benchmark.
The bug localization problem can be formulated as follows: given an issue with a bug description and a repository snapshot in a state where the bug is reproducible, identify the files within the repository that need to be modified to address the reported bug.
The dataset provides all the required components for evaluation of bug localization… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-bug-localization.climbmix-40b-az
ClimbMix 40B — Azerbaijani
A large-scale Azerbaijani text dataset created by translating the English karpathy/climbmix-400b-shuffle dataset into Azerbaijani using Google Translate.
Dataset Summary
This dataset contains approximately 40 billion tokens of Azerbaijani text, making it one of the largest publicly available Azerbaijani language corpora. It is intended for pretraining and fine-tuning large language models (LLMs) for the Azerbaijani language.
Property
Value… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/climbmix-40b-az.local-llm-benchmark
Local LLM Benchmark — Technical and Uncensored Behavior (NVIDIA RTX 5070 Ti 16GB)
English | 简体中文 | 繁體中文 | 한국어 | Español | 日本語 | हिन्दी | Русский | Português | తెలుగు | Français | Deutsch | Italiano | Tiếng Việt | العربية | اردو | বাংলা | فارسی | Română | Türkçe
Manual evaluation results of local GGUF model variants on a single consumer machine,
combining two fully independent benchmarks:
technical/
uncensored/
Measures
capability: coding, systems, networking, DB, agents… See the full description on the dataset page: https://huggingface.co/datasets/nanimani/local-llm-benchmark.azerbaijani-pretrain-corpus
Azerbaijani Pretraining Corpus (merged & deduplicated)
A cleaned Azerbaijani text corpus assembled for language-model pretraining,
merging two curated sources and removing exact duplicates.
Contents
Documents: 6,931,898
Tokens: ~5.36B (measured with the o200k_base tokenizer; an
Azerbaijani-specific tokenizer will yield fewer tokens, as o200k_base
segments agglutinative Azerbaijani inefficiently)
Avg tokens/document: ~773
Fields
text — the… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-pretrain-corpus.benchname-bug-localization
🥷 BenchName (Bug localization)
This is the benchmark for the Bug localization task as part of the
🥷 BenchName benchmark.
The bug localization problem can be formulated as follows: given an issue with a bug description and a repository snapshot in a state where the bug is reproducible, identify the files within the repository that need to be modified to address the reported bug.
The dataset provides all the required components for evaluation of bug localization approaches in… See the full description on the dataset page: https://huggingface.co/datasets/anon-iclr-submission/benchname-bug-localization.AzTC
AzTC (Azerbaijan Text Corpus)
This is the first version of the largest text corpus in the Azerbaijani language.
Overview
The AzTC contains 51 million (approximately 1 billion tokens) non-recurring sentences.
The data was collected from various resources such as websites, news, books, wikipedia, legislation, scientific articles and etc.
License
The AzTC licensed under the CC BY-NC-ND 4.0 license.
What does this license allow?
Attribution: You must give… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/AzTC.Bilik-Instruct
Bilik-Instruct: Azerbaijani Persona-Driven SFT Dataset
Bilik-Instruct is a large-scale, high-quality Supervised Fine-Tuning (SFT) dataset for the Azerbaijani language. It is built upon the LocalDoc/wikipedia_azerbaijan dataset and enhanced using a novel persona-driven generation technique via OpenAI's GPT-5.
The goal of this dataset is to move beyond formal, encyclopedic language and capture natural, conversational Azerbaijani across various domains.
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/Bilik-Instruct.selective-learning-benchmark-ip
Selective Learning Benchmark Data: Inoculation Prompting
This repository is an inoculation-prompting variant of localized-ft/selective-learning-benchmark. It bundles selective-learning task data in task_data_model_v1 JSONL format and prepends a subset-specific inoculation prompt as the system turn of every sft and validation example. The eval and control examples intentionally omit the prompt so evaluation measures learned behavior rather than direct prompt steering.
Each task… See the full description on the dataset page: https://huggingface.co/datasets/localized-ft/selective-learning-benchmark-ip.thai-local-instruction-v2
Thai local instruction v2
Thai local language instruction dataset v2
List:
korat (ภาษาโคราช)
pattani (ภาษาปักษ์ใต้หรือภาษาใต้)
khummuang (ภาษาเหนือหรือภาษาคำเมือง)
isan (ภาษาอีสาน)
Sources:
th.wiktionary.org (CC BY-SA) for khummuang dictionary.
isan.clubs.chula.ac.th (CC BY-SA-NC) for isan dictionary and sentence.
pythainlp/thai-local-language-translation-dataset (CC BY-SA) for korat sentences, pattani sentences, and khummuang sentences.
Created by Wannaphong Phatthiyaphaibun
AzTC-full
AzTC — Full Version (Azerbaijan Text Corpus)
The expanded version of LocalDoc/AzTC,
and one of the largest text corpora in the Azerbaijani language.
Overview
The corpus contains approximately 2.4 billion tokens of Azerbaijani text,
compiled and cleaned from a wide range of sources including news portals, books,
Wikipedia, and legislation. Text is organized at the document / passage level so
that each row is a coherent unit rather than an isolated fragment.… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/AzTC-full.local-agentic-coding-bench-8gb-vram-2026-05
agentic coding benchmark: local LLMs on 8GB VRAM
can local LLMs do agentic coding (multi-turn tool calling, file creation, debugging) on consumer hardware? this dataset captures real test results.
hardware
GPU: NVIDIA RTX 4060 Ti 8GB
CPU: Intel i7-14700F
RAM: 32 GB DDR5
OS: Windows 11 + WSL2 (Ubuntu)
inference: llama-server (turboquant fork of llama.cpp)
what was tested
two agent frameworks:
Hermes Agent (NousResearch): structured tool calling with… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/local-agentic-coding-bench-8gb-vram-2026-05.community_oscar_azerbaijani
Community-OSCAR Azerbaijani
This is Azerbaijani version Community OSCAR dataset https://huggingface.co/datasets/oscar-corpus/community-oscar.
Dataset Statistics (Aggregate)
Metric
Value
Language
Azerbaijani (az)
Average per release
3.36 GiB, 603,832 documents
Words per release
~408.8M words
Characters per release
~3.12B characters
Total size (all releases)
137.62 GiB
Total lines
24.76M
Total words
16.76B words
Total characters
128.07B characters… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/community_oscar_azerbaijani.thai-sent-local-v2
Thai sent local v2
List:
korat (ภาษาโคราช)
pattani (ภาษาปักษ์ใต้หรือภาษาใต้)
khummuang (ภาษาเหนือหรือภาษาคำเมือง)
isan (ภาษาอีสาน)
Sources:
th.wiktionary.org (CC BY-SA) for khummuang dictionary.
isan.clubs.chula.ac.th (CC BY-SA-NC) for isan dictionary and sentence.
pythainlp/thai-local-language-translation-dataset (CC BY-SA) for korat sentences, pattani sentences, and khummuang sentences.
Created by Wannaphong Phatthiyaphaibun
selective-learning-benchmark
Selective Learning Benchmark Data
This repository bundles selective-learning task data from Sunday, Srija, and Sultan in task_data_model_v1 JSONL format.
Each task directory contains a manifest.json with contributor/source attribution, a capability description, an unintended-generalization description, split files, and row counts.
Each Hugging Face config/subset is one dataset named as [type]-[name], with sft, validation, eval, and control splits where available. The type values… See the full description on the dataset page: https://huggingface.co/datasets/localized-ft/selective-learning-benchmark.news_azerbaijan_2Azerbaijani News Dataset
Description
This dataset contains news from https://musavat.com/ in Azerbaijani language. It was created in 2024 and contains 753k news (approximately 11 million sentences).
Format
The dataset is provided in comma-separated values (CSV) format. Each article is represented on a new line with the following fields separated by commas:
id: news unique id
date: news date
category: news category
title: news title
text: news text
License
Copyright of the content belongs to… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/news_azerbaijan_2.books_datasetAzerbaijani Books Dataset
Description
This dataset contains 2800 books on different topics in Azerbaijani language. It was created in 2024 and contains 7.8 million sentences.
The books were divided into sentences and pre-filtered.
The dataset included only those sentences where the percentage of letters was at least 80% of the total number of characters.
The sequence of sentences is the same as in books.
Format
The dataset is provided in comma-separated values (CSV) format. Each article is… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/books_dataset.community_oscar_azerbaijani_scored
Azerbaijani Web Corpus with Quality Scores
This dataset is the full Azerbaijani web corpus
LocalDoc/community_oscar_azerbaijani
with a continuous quality score attached to every document. It is intended
as the filtering layer for building a clean Azerbaijani pretraining corpus:
each document carries a score that lets you keep, clean, or drop it according
to your own thresholds.
What was done
Every document in the source corpus was scored by the model… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/community_oscar_azerbaijani_scored.news_azerbaijanAzerbaijani News Dataset
Description
This dataset contains news from https://axar.az in Azerbaijani language. It was created in 2024 and contains 447k news.
Format
The dataset is provided in comma-separated values (CSV) format. Each article is represented on a new line with the following fields separated by commas:
date: news date
id: news unique id
title: news title
text: news text
License
Copyright of the content belongs to https://axar.az resource. Citation is mandatory when using… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/news_azerbaijan.LocalLLaMA-posts
r/LocalLLaMA posts
Posts from r/LocalLLaMA pulled up through Tue Mar 3 9PM EST 2026 with arctic-shift. Now you can check if your wonderfully thought out post hasn't already been asked 30x
Usage
For simple semantic search, try loading it in the vectorsearch-hub-datasets space:
mlx-local-inference-benchmarks
MLX local-inference benchmarks — Qwen3.6 & Laguna-S/XS families
Raw results, harnesses and methodology for an 8-axis benchmark of four MLX
checkpoints on a 128 GB M5 Max. Everything a person would need to check my numbers
or disagree with them.
Companion model repos:
Tess-4-27B-MLX-Q8 — with a working MTP head
Tess-4-27B-MLX-Q4 — same, at 4-bit
NEW (2026-07-24): the Laguna chapter — REPORT-LAGUNA.md + results-laguna/
Five-way same-engine bake-off (Laguna-S… See the full description on the dataset page: https://huggingface.co/datasets/studioburnside/mlx-local-inference-benchmarks.wikipedia_azerbaijanAzerbaijani Wikipedia Dataset
Description
This dataset contains all articles from Wikipedia in Azerbaijani language. It was created in 2024 and contains 260k articles.
Format
The dataset is provided in comma-separated values (CSV) format. Each article is represented on a new line with the following fields separated by commas:
title: Title of the article
text: Text of the article
url: URL of the article
License
The dataset is licensed under the Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/wikipedia_azerbaijan.localagent-dispatch-data
LocalAgent Dispatch Data
Synthetic data for training/evaluating a generable tool-dispatch model over a 50-tool surface
(route head → dense selector → pointer-copy). A static snapshot of the deterministic generators in
LocalAgent (src/localagent/data/). Train/eval are
disjoint in both phrasing and slot values. Companion model + demo:
danelcsb/localagent-tiny-30m-byte ·
Space.
Configs
config
rows (train/eval)
what it is
paraphrase
1000 / 1000
many natural… See the full description on the dataset page: https://huggingface.co/datasets/danelcsb/localagent-dispatch-data.LocalLLaMA-comments
LocalLLaMA-comments
A companion dataset to pszemraj/LocalLLaMA-posts. Time frame is in sync (up through Tue Mar 3 9PM EST 2026)
smolified-bengali-local-food-guide
🤏 smolified-bengali-local-food-guide
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-bengali-local-food-guide.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 638d3b25)
Records: 1050
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
local-code-arena-deepseek-r1_1.5b
Local Code Arena Telemetry: MBPP Benchmark on DeepSeek R1 1.5B
This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the DeepSeek R1 1.5B distilled reasoning architecture.
This specific run establishes the performance boundaries of lightweight reasoning models under strict execution time limits on consumer hardware.
📊 Core Performance… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-deepseek-r1_1.5b.azerbaijan-history-reasoning-SFTThis dataset was developed based on https://huggingface.co/datasets/LocalDoc/azerbaijan_history_corpus
Factual: Simple questions that ask for specific facts from one part of the text.
Synthesis: Questions that require summarizing, comparing, or logically inferring from multiple different parts of the text.
General/Contextual: General questions inspired by the text, either without direct reference to the context or with a reference like "Based on this text...".
local-code-master_telemetry_arena
Local Code Arena: Comprehensive Telemetry Matrix Dataset
🏆 An Empirical Dataset tracking Local Generation Throughput (TPS), Real-Time Latency, Syntactic CodeBLEU Alignments, and Functional Pass Rates across 22 Edge Architectures.
📊 Dataset Blueprint
This dataset contains a consolidated, high-fidelity matrix of 11,000 unique token-generation execution loops across 22 state-of-the-art open-weights language models (ranging from 500M to 15.5B parameters). Every… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-master_telemetry_arena.numina_math_azerbaijaniThis is part of the translated version of the original dataset: https://huggingface.co/datasets/AI-MO/NuminaMath-CoT
