datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ASearcher-Local-Knowledgetyped-decisions
Typed Decisions
A benchmark for typed probabilistic decisions over shared state. You give a
model one piece of unstructured state. It answers several typed questions about
that state at once, and every answer is a probability distribution rather than a
single label.
The schema follows the System One primitives used by
TypeSafe AI: noul, choice and score. A row replays against any API that implements that shape. This benchmark is
independent. It is not affiliated with TypeSafe… See the full description on the dataset page: https://huggingface.co/datasets/LocalLLaMA/typed-decisions.aidovecl-vehicle-detection-classification-localization
AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization
We introduce an annotated AI-generated dataset of eye-level vehicle images using outpainting, offering versatile generation of diverse vehicle classes in varied contexts with pretrained models.
Citation Notice
Please ensure that all publications and presentations using this data reference the following paper:
Kazemi, A., Fatima, Q. ul A., Kindratenko, V., & Tessum, C. W.… See the full description on the dataset page: https://huggingface.co/datasets/amir-kazemi/aidovecl-vehicle-detection-classification-localization.terminal-bench-traces-localis25-local-distortionslca-bug-localization
🏟️ Long Code Arena (Bug localization)
This is the benchmark for the Bug localization task as part of the
🏟️ Long Code Arena benchmark.
The bug localization problem can be formulated as follows: given an issue with a bug description and a repository snapshot in a state where the bug is reproducible, identify the files within the repository that need to be modified to address the reported bug.
The dataset provides all the required components for evaluation of bug localization… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-bug-localization.swebench-localisation
Finding the file: localisation on SWE-bench Verified
Given a GitHub issue, which file do you have to change? This is the retrieval step every coding
agent performs before it writes a patch, and none of the leaderboards score it separately. SWE-bench's five
leaderboards all score % Resolved, which folds localisation and patch-writing into one
number.
This bundle is that step measured on its own, on all 500 instances of SWE-bench Verified, with a
floor. The write-up is
Finding… See the full description on the dataset page: https://huggingface.co/datasets/RiverRider/swebench-localisation.split-avelina-python-edummtrailer-pe-av-unimodal-local-embeddingsazerbaijani_asr
Azerbaijani ASR Dataset
Dataset Description
This dataset contains Azerbaijani speech data for Automatic Speech Recognition (ASR) tasks.
Dataset Summary
Language: Azerbaijani (az)
Task: Automatic Speech Recognition
Total Duration: ~328 hours
Total Samples: ~345,643 audio-text pairs
Audio Format: WAV, 16kHz sampling rate
License: CC-BY-4.0
Dataset Structure
Each audio segment is specially numbered so that you can merge them if you… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_asr.climbmix-40b-az
ClimbMix 40B — Azerbaijani
A large-scale Azerbaijani text dataset created by translating the English karpathy/climbmix-400b-shuffle dataset into Azerbaijani using Google Translate.
Dataset Summary
This dataset contains approximately 40 billion tokens of Azerbaijani text, making it one of the largest publicly available Azerbaijani language corpora. It is intended for pretraining and fine-tuning large language models (LLMs) for the Azerbaijani language.
Property
Value… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/climbmix-40b-az.localizer
Dartbrains Localizer Dataset
A subset of the Brainomics/Localizer functional MRI dataset, prepared for the Dartbrains neuroimaging course at Dartmouth College.
Quick Start
Load beta maps (recommended for most exercises)
from datasets import load_dataset
ds = load_dataset("dartbrains/localizer", "betas")
img = ds[0]["nifti"] # nibabel.Nifti1Image
subject = ds[0]["subject"] # "S01"
condition = ds[0]["condition"] # "audio_computation"… See the full description on the dataset page: https://huggingface.co/datasets/dartbrains/localizer.split-finemathlocale-benchmark-sra500
Locale Embedding Benchmark : SRA 500
What
This benchmark contains embeddings of raw genomic read sequences produced by the LOCALE DNA transformer model to test the use of vector search over large sequence repositories like the NIH Sequence Read Archive. The benchmark contains the embeddings of 163,578,486 sequence embeddings coming from 500 SRA Accessions. All vectors are Float32, D=768, roughly 500GB of total data.
Given a read, we want to find accessions… See the full description on the dataset page: https://huggingface.co/datasets/rsynk/locale-benchmark-sra500.yt-pe-av-unimodal-local-embeddingscleangov-local-settlements
지방재정365 결산 통계 Open API (세입·세출결산, 재무제표, 지방세, 지역통합재정통계, 공공시설·청사·채무)
지방재정365 재정데이터개방 허브의 "결산" 분류 33 서비스. 세출결산(기능별·성질별·회계별·구조별 단체별), 세입결산(재원별·성질별), 기금결산, 교육비특별회계 결산, 투자적경비 순계, 재무제표(재정상태표·통합재정운영표·순자산변동표·복식부기 수익·비용·자산·부채), 지방세(징수율·세목별 비중·세수신장률·체납 누계), 지역통합재정통계 (세입·세출·자산·부채·인건비·업무추진비·행사경비 비율), 공공시설운영현황, 청사면적, 채무현황. 자치단체 재정의 결산 기준 정본이며 FISCAL-LOC-002(세부사업별 세출 XLSX) 보다 집계 수준이 높고 분류 축이 다양하다. 비교군·기관 개요 화면의 결산 수치를 여기서 낸다.
출처: https://www.lofin365.go.kr/portal/LF5100000.do
이용 조건: 허브 명세 이용조건… See the full description on the dataset page: https://huggingface.co/datasets/eddmpython/cleangov-local-settlements.deepswe-mini
deepswe-mini
16 of the 113 tasks in DeepSWE v1.1, picked so that running just these ranks models the same way the full benchmark does.
DeepSWE is a good benchmark and an expensive one. Every task is a long-horizon feature request in its own container, and a full pass takes close to two days of agent time run one task at a time. If you are comparing models, agent harnesses or prompts, and the differences you care about are more than a few points, these 16 tasks give you the same… See the full description on the dataset page: https://huggingface.co/datasets/LocalLLaMA/deepswe-mini.LOCUS-v1
LOCUS v1.0
This repository contains the dataset presented in the paper Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States.
Dataset Summary
LOCUS v1.0 is a chunk-level dataset of U.S. municipal and county law text labeled by legal function. Each eligible chunk is assigned a function, a binary is_substantive label, and all substantive provisions are assigned a topic.
The dataset is intended for legal text research, local-law structure… See the full description on the dataset page: https://huggingface.co/datasets/LocalLaws/LOCUS-v1.terminal-bench-mini
terminal-bench-mini
Fourteen of Terminal-Bench 2.0's ninety tasks, picked so that ranking agents on
the subset reproduces ranking them on the whole benchmark.
Running ninety tasks five times each is how the official leaderboard is built.
That is out of reach if you are comparing quant variants, fine-tunes or local
models on your own hardware. This subset turns a multi-day sweep into a few
hours.
Same approach as deepswe-mini:
take the published per-task results, rank the field… See the full description on the dataset page: https://huggingface.co/datasets/LocalLLaMA/terminal-bench-mini.local-llm-benchmark
Local LLM Benchmark — Technical and Uncensored Behavior (NVIDIA RTX 5070 Ti 16GB)
English | 简体中文 | 繁體中文 | 한국어 | Español | 日本語 | हिन्दी | Русский | Português | తెలుగు | Français | Deutsch | Italiano | Tiếng Việt | العربية | اردو | বাংলা | فارسی | Română | Türkçe
Manual evaluation results of local GGUF model variants on a single consumer machine,
combining two fully independent benchmarks:
technical/
uncensored/
Measures
capability: coding, systems, networking, DB, agents… See the full description on the dataset page: https://huggingface.co/datasets/nanimani/local-llm-benchmark.movielens-pe-av-local-embeddingslocal_administrations_directory-full-documents
🇫🇷 Référentiel des administrations locales – Version structurée
Ce dataset regroupe l’Annuaire de l’administration – Base de données locales, qui recense l’ensemble des administrations et services publics locaux français :
collectivités territoriales,
services municipaux,
services départementaux et régionaux,
établissements publics locaux,
structures administratives de proximité.
Les données sont issues des sources open data officielles publiées sur data.gouv.fr et… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/local_administrations_directory-full-documents.local-administrations-directory
📢 Sondage 2026 : Utilisation des datasets publiques de MediaTech
Vous utilisez ce dataset ou d’autres datasets de notre collection MediaTech ? Votre avis compte !
Aidez-nous à améliorer nos datasets publiques en répondant à ce sondage rapide (5 min) : 👉 https://grist.numerique.gouv.fr/o/albert/forms/gF4hLaq9VvUog6c5aVDuMw/11
Merci pour votre contribution ! 🙌
🇫🇷 French Local Administrations Directory Dataset
This dataset is a processed and embedded version… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/local-administrations-directory.YOXLA-Benchmark
YOXLA Benchmark
1443 frozen examples for evaluating large language models in
Azerbaijani, across four blocks and eleven tasks.
Run with the YOXLA framework:
pip install "yoxla[api]"
yoxla run --provider openrouter --model <model> --block all
Or load a config directly:
from datasets import load_dataset
data = load_dataset("LocalDoc/YOXLA-Benchmark", "rag_selection_v1")["test"]
What makes it different
Every answer space is closed. A label, a number, or a span… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/YOXLA-Benchmark.azerbaijani-pretrain-corpus
Azerbaijani Pretraining Corpus (merged & deduplicated)
A cleaned Azerbaijani text corpus assembled for language-model pretraining,
merging two curated sources and removing exact duplicates.
Contents
Documents: 6,931,898
Tokens: ~5.36B (measured with the o200k_base tokenizer; an
Azerbaijani-specific tokenizer will yield fewer tokens, as o200k_base
segments agglutinative Azerbaijani inefficiently)
Avg tokens/document: ~773
Fields
text — the… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-pretrain-corpus.bug-localization
Bug Localization
This is the data for Bug Localization benchmark.
How-to
Since the dataset is private, if you haven't used HF Hub before, add your token via huggingface-cli first:
huggingface-cli login
List all the available configs via datasets.get_dataset_config_names and choose an appropriate one
Load the data via load_dataset:
from datasets import load_dataset
# Select a configuration from ["py", "java", "kt", "mixed"]
configuration = "py"
# Select a split from… See the full description on the dataset page: https://huggingface.co/datasets/tiginamaria/bug-localization.cleangov-local-fiscal-metrics
지방재정365 통합공시·재정분석 Open API (수의계약·행사축제·업무추진비·의회경비·채무·기금·지방보조금·재정분석 결과)
지방재정365 재정데이터개방 허브의 "지방재정 통합공시" 분류 77 서비스와 "성과/평가 > 재정분석결과" 5 서비스, 합계 82 서비스. 자치단체별로 공시하는 지표군 (수의계약비율, 행사·축제경비 비율·편성내역·원가회계, 업무추진비 비율·절감률·기관운영·시책추진, 지방의회 관련경비·국외여비, 채무·지방채·보증채무·채권, 기금현재액, 지방보조금(2016·2020·2021)·지방보조금비율, 재정자립도·재정자주도·통합재정수지(결산·최종), 사회보장적 수혜금, 예비비, 통장이장반장 보상금, 공무원 관련경비, 지방교부세 인센티브·자체노력 반영, 재정운용계획 등) 과 행정안전부 지방재정분석 결과 지표다. 동종 지자체 비교(비교군) 의 기준 자료이며 대부분 자치단체 × 회계연도 × 지표 단위의 집계값이라 사건 단위 연결에는 제한이 있다.… See the full description on the dataset page: https://huggingface.co/datasets/eddmpython/cleangov-local-fiscal-metrics.azerbaijani_retriever_corpus
A Large-Scale Azerbaijani Corpus for Contrastive Retriever Training
Dataset Description
This dataset is a large-scale, high-quality resource designed for training Azerbaijani text embedding models for information retrieval tasks. It contains 671,528 training instances, each consisting of a query, a relevant positive document, and 10 hard-negative documents.
The primary goal of this dataset is to facilitate the training of dense retriever models using contrastive learning.… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_retriever_corpus.cleangov-local-budgets
지방재정365 예산 통계·재정지표 Open API
지방재정365 예산 분류의 공식 통계다. 분야·회계·재원별 예산 구성, 교육·인력·의회 관련 경비, 주민 1인당 금액과 사업 비중을 제공한다. 예산 계획을 설명하는 자료이며 실제 지급액이나 결산액으로 해석하지 않는다. 서비스별 제공 항목과 단위는 resources가 소유한다. 재정자립도·재정자주도·통합재정수지비율은 FISCAL-LOC-003의 통합공시에서 제공한다.
출처: https://www.lofin365.go.kr/portal/LF5100000.do
이용 조건: 허브 명세 이용조건 "출처표시, 상업적·비상업적 이용가능, 변형 등 2차적 저작물 작성 가능" (공공누리 1유형 상당)
manifest.json은 현재 검증된 파일의 계층·원래 경로·SHA-256·크기·불변 객체 경로를 제공한다. 같은 commit revision으로 manifest와 객체를 내려받고 해시를 검증한다. 예산현액과 지출액은… See the full description on the dataset page: https://huggingface.co/datasets/eddmpython/cleangov-local-budgets.yt-pe-av-local-embeddings
