CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01inclusionAI /ASearcher-Local-Knowledgetext10M<n<100M7 likes15k downloads1y agoHugging Face02LocalLLaMA /typed-decisions Typed Decisions A benchmark for typed probabilistic decisions over shared state. You give a model one piece of unstructured state. It answers several typed questions about that state at once, and every answer is a probability distribution rather than a single label. The schema follows the System One primitives used by TypeSafe AI: noul, choice and score. A row replays against any API that implements that shape. This benchmark is independent. It is not affiliated with TypeSafe… See the full description on the dataset page: https://huggingface.co/datasets/LocalLLaMA/typed-decisions.tabulartext-classification1K<n<10K16 likes5k downloads17h agoHugging Face03amir-kazemi /aidovecl-vehicle-detection-classification-localization AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization We introduce an annotated AI-generated dataset of eye-level vehicle images using outpainting, offering versatile generation of diverse vehicle classes in varied contexts with pretrained models. Citation Notice Please ensure that all publications and presentations using this data reference the following paper: Kazemi, A., Fatima, Q. ul A., Kindratenko, V., & Tessum, C. W.… See the full description on the dataset page: https://huggingface.co/datasets/amir-kazemi/aidovecl-vehicle-detection-classification-localization.imageobject-detection1K<n<10K0 likes4k downloads5mo agoHugging Face04mlfoundations-dev /terminal-bench-traces-localtext1K<n<10K0 likes1.8k downloads1y agoHugging Face05kuhlmannm /is25-local-distortionsaudio1K<n<10K0 likes1.6k downloads1y agoHugging Face06JetBrains-Research /lca-bug-localization 🏟️ Long Code Arena (Bug localization) This is the benchmark for the Bug localization task as part of the 🏟️ Long Code Arena benchmark. The bug localization problem can be formulated as follows: given an issue with a bug description and a repository snapshot in a state where the bug is reproducible, identify the files within the repository that need to be modified to address the reported bug. The dataset provides all the required components for evaluation of bug localization… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-bug-localization.imagetext-generation10K<n<100K4 likes1.2k downloads2y agoHugging Face07RiverRider /swebench-localisation Finding the file: localisation on SWE-bench Verified Given a GitHub issue, which file do you have to change? This is the retrieval step every coding agent performs before it writes a patch, and none of the leaderboards score it separately. SWE-bench's five leaderboards all score % Resolved, which folds localisation and patch-writing into one number. This bundle is that step measured on its own, on all 500 instances of SWE-bench Verified, with a floor. The write-up is Finding… See the full description on the dataset page: https://huggingface.co/datasets/RiverRider/swebench-localisation.tabulartext-retrievaln<1K3 likes1.1k downloads1d agoHugging Face08LocalResearchGroup /split-avelina-python-edutabular1M<n<10M0 likes893 downloads1y agoHugging Face09seungheondoh /mmtrailer-pe-av-unimodal-local-embeddingstext10K<n<100K0 likes794 downloads6mo agoHugging Face10LocalDoc /azerbaijani_asr Azerbaijani ASR Dataset Dataset Description This dataset contains Azerbaijani speech data for Automatic Speech Recognition (ASR) tasks. Dataset Summary Language: Azerbaijani (az) Task: Automatic Speech Recognition Total Duration: ~328 hours Total Samples: ~345,643 audio-text pairs Audio Format: WAV, 16kHz sampling rate License: CC-BY-4.0 Dataset Structure Each audio segment is specially numbered so that you can merge them if you… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_asr.audioautomatic-speech-recognition100K<n<1M4 likes669 downloads2mo agoHugging Face11LocalDoc /climbmix-40b-az ClimbMix 40B — Azerbaijani A large-scale Azerbaijani text dataset created by translating the English karpathy/climbmix-400b-shuffle dataset into Azerbaijani using Google Translate. Dataset Summary This dataset contains approximately 40 billion tokens of Azerbaijani text, making it one of the largest publicly available Azerbaijani language corpora. It is intended for pretraining and fine-tuning large language models (LLMs) for the Azerbaijani language. Property Value… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/climbmix-40b-az.texttext-generation10M<n<100M1 likes647 downloads6mo agoHugging Face12dartbrains /localizer Dartbrains Localizer Dataset A subset of the Brainomics/Localizer functional MRI dataset, prepared for the Dartbrains neuroimaging course at Dartmouth College. Quick Start Load beta maps (recommended for most exercises) from datasets import load_dataset ds = load_dataset("dartbrains/localizer", "betas") img = ds[0]["nifti"] # nibabel.Nifti1Image subject = ds[0]["subject"] # "S01" condition = ds[0]["condition"] # "audio_computation"… See the full description on the dataset page: https://huggingface.co/datasets/dartbrains/localizer.imageimage-classificationn<1K1 likes620 downloads3mo agoHugging Face13LocalResearchGroup /split-finemathtabular1M<n<10M0 likes587 downloads1y agoHugging Face14rsynk /locale-benchmark-sra500 Locale Embedding Benchmark : SRA 500 What This benchmark contains embeddings of raw genomic read sequences produced by the LOCALE DNA transformer model to test the use of vector search over large sequence repositories like the NIH Sequence Read Archive. The benchmark contains the embeddings of 163,578,486 sequence embeddings coming from 500 SRA Accessions. All vectors are Float32, D=768, roughly 500GB of total data. Given a read, we want to find accessions… See the full description on the dataset page: https://huggingface.co/datasets/rsynk/locale-benchmark-sra500.textfeature-extractionn<1K0 likes583 downloads2d agoHugging Face15seungheondoh /yt-pe-av-unimodal-local-embeddingstext10K<n<100K0 likes510 downloads5mo agoHugging Face16eddmpython /cleangov-local-settlements 지방재정365 결산 통계 Open API (세입·세출결산, 재무제표, 지방세, 지역통합재정통계, 공공시설·청사·채무) 지방재정365 재정데이터개방 허브의 "결산" 분류 33 서비스. 세출결산(기능별·성질별·회계별·구조별 단체별), 세입결산(재원별·성질별), 기금결산, 교육비특별회계 결산, 투자적경비 순계, 재무제표(재정상태표·통합재정운영표·순자산변동표·복식부기 수익·비용·자산·부채), 지방세(징수율·세목별 비중·세수신장률·체납 누계), 지역통합재정통계 (세입·세출·자산·부채·인건비·업무추진비·행사경비 비율), 공공시설운영현황, 청사면적, 채무현황. 자치단체 재정의 결산 기준 정본이며 FISCAL-LOC-002(세부사업별 세출 XLSX) 보다 집계 수준이 높고 분류 축이 다양하다. 비교군·기관 개요 화면의 결산 수치를 여기서 낸다. 출처: https://www.lofin365.go.kr/portal/LF5100000.do 이용 조건: 허브 명세 이용조건… See the full description on the dataset page: https://huggingface.co/datasets/eddmpython/cleangov-local-settlements.text1M<n<10M0 likes503 downloads16d agoHugging Face17LocalLLaMA /deepswe-mini deepswe-mini 16 of the 113 tasks in DeepSWE v1.1, picked so that running just these ranks models the same way the full benchmark does. DeepSWE is a good benchmark and an expensive one. Every task is a long-horizon feature request in its own container, and a full pass takes close to two days of agent time run one task at a time. If you are comparing models, agent harnesses or prompts, and the differences you care about are more than a few points, these 16 tasks give you the same… See the full description on the dataset page: https://huggingface.co/datasets/LocalLLaMA/deepswe-mini.tabularn<1K4 likes497 downloads6d agoHugging Face18LocalLaws /LOCUS-v1 LOCUS v1.0 This repository contains the dataset presented in the paper Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States. Dataset Summary LOCUS v1.0 is a chunk-level dataset of U.S. municipal and county law text labeled by legal function. Each eligible chunk is assigned a function, a binary is_substantive label, and all substantive provisions are assigned a topic. The dataset is intended for legal text research, local-law structure… See the full description on the dataset page: https://huggingface.co/datasets/LocalLaws/LOCUS-v1.tabulartext-classification1M<n<10M93 likes447 downloads3mo agoHugging Face19LocalLLaMA /terminal-bench-mini terminal-bench-mini Fourteen of Terminal-Bench 2.0's ninety tasks, picked so that ranking agents on the subset reproduces ranking them on the whole benchmark. Running ninety tasks five times each is how the official leaderboard is built. That is out of reach if you are comparing quant variants, fine-tunes or local models on your own hardware. This subset turns a multi-day sweep into a few hours. Same approach as deepswe-mini: take the published per-task results, rank the field… See the full description on the dataset page: https://huggingface.co/datasets/LocalLLaMA/terminal-bench-mini.tabularn<1K3 likes446 downloads3d agoHugging Face20nanimani /local-llm-benchmark Local LLM Benchmark — Technical and Uncensored Behavior (NVIDIA RTX 5070 Ti 16GB) English | 简体中文 | 繁體中文 | 한국어 | Español | 日本語 | हिन्दी | Русский | Português | తెలుగు | Français | Deutsch | Italiano | Tiếng Việt | العربية | اردو | বাংলা | فارسی | Română | Türkçe Manual evaluation results of local GGUF model variants on a single consumer machine, combining two fully independent benchmarks: technical/ uncensored/ Measures capability: coding, systems, networking, DB, agents… See the full description on the dataset page: https://huggingface.co/datasets/nanimani/local-llm-benchmark.tabulartext-generation1K<n<10K2 likes445 downloads7d agoHugging Face21seungheondoh /movielens-pe-av-local-embeddingstext10K<n<100K0 likes408 downloads6mo agoHugging Face22hulk10 /local_administrations_directory-full-documents 🇫🇷 Référentiel des administrations locales – Version structurée Ce dataset regroupe l’Annuaire de l’administration – Base de données locales, qui recense l’ensemble des administrations et services publics locaux français : collectivités territoriales, services municipaux, services départementaux et régionaux, établissements publics locaux, structures administratives de proximité. Les données sont issues des sources open data officielles publiées sur data.gouv.fr et… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/local_administrations_directory-full-documents.text10K<n<100K1 likes395 downloads21h agoHugging Face23AgentPublic /local-administrations-directory 📢 Sondage 2026 : Utilisation des datasets publiques de MediaTech Vous utilisez ce dataset ou d’autres datasets de notre collection MediaTech ? Votre avis compte ! Aidez-nous à améliorer nos datasets publiques en répondant à ce sondage rapide (5 min) : 👉 https://grist.numerique.gouv.fr/o/albert/forms/gF4hLaq9VvUog6c5aVDuMw/11 Merci pour votre contribution ! 🙌 🇫🇷 French Local Administrations Directory Dataset This dataset is a processed and embedded version… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/local-administrations-directory.text10K<n<100K1 likes394 downloads11d agoHugging Face24LocalDoc /YOXLA-Benchmark YOXLA Benchmark 1443 frozen examples for evaluating large language models in Azerbaijani, across four blocks and eleven tasks. Run with the YOXLA framework: pip install "yoxla[api]" yoxla run --provider openrouter --model <model> --block all Or load a config directly: from datasets import load_dataset data = load_dataset("LocalDoc/YOXLA-Benchmark", "rag_selection_v1")["test"] What makes it different Every answer space is closed. A label, a number, or a span… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/YOXLA-Benchmark.tabularquestion-answering1K<n<10K1 likes380 downloads10d agoHugging Face25LocalDoc /azerbaijani-pretrain-corpus Azerbaijani Pretraining Corpus (merged & deduplicated) A cleaned Azerbaijani text corpus assembled for language-model pretraining, merging two curated sources and removing exact duplicates. Contents Documents: 6,931,898 Tokens: ~5.36B (measured with the o200k_base tokenizer; an Azerbaijani-specific tokenizer will yield fewer tokens, as o200k_base segments agglutinative Azerbaijani inefficiently) Avg tokens/document: ~773 Fields text — the… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-pretrain-corpus.texttext-generation1M<n<10M0 likes378 downloads4mo agoHugging Face26tiginamaria /bug-localization Bug Localization This is the data for Bug Localization benchmark. How-to Since the dataset is private, if you haven't used HF Hub before, add your token via huggingface-cli first: huggingface-cli login List all the available configs via datasets.get_dataset_config_names and choose an appropriate one Load the data via load_dataset: from datasets import load_dataset # Select a configuration from ["py", "java", "kt", "mixed"] configuration = "py" # Select a split from… See the full description on the dataset page: https://huggingface.co/datasets/tiginamaria/bug-localization.tabular10K<n<100K3 likes339 downloads2y agoHugging Face27eddmpython /cleangov-local-fiscal-metrics 지방재정365 통합공시·재정분석 Open API (수의계약·행사축제·업무추진비·의회경비·채무·기금·지방보조금·재정분석 결과) 지방재정365 재정데이터개방 허브의 "지방재정 통합공시" 분류 77 서비스와 "성과/평가 > 재정분석결과" 5 서비스, 합계 82 서비스. 자치단체별로 공시하는 지표군 (수의계약비율, 행사·축제경비 비율·편성내역·원가회계, 업무추진비 비율·절감률·기관운영·시책추진, 지방의회 관련경비·국외여비, 채무·지방채·보증채무·채권, 기금현재액, 지방보조금(2016·2020·2021)·지방보조금비율, 재정자립도·재정자주도·통합재정수지(결산·최종), 사회보장적 수혜금, 예비비, 통장이장반장 보상금, 공무원 관련경비, 지방교부세 인센티브·자체노력 반영, 재정운용계획 등) 과 행정안전부 지방재정분석 결과 지표다. 동종 지자체 비교(비교군) 의 기준 자료이며 대부분 자치단체 × 회계연도 × 지표 단위의 집계값이라 사건 단위 연결에는 제한이 있다.… See the full description on the dataset page: https://huggingface.co/datasets/eddmpython/cleangov-local-fiscal-metrics.text100K<n<1M0 likes328 downloads15d agoHugging Face28LocalDoc /azerbaijani_retriever_corpus A Large-Scale Azerbaijani Corpus for Contrastive Retriever Training Dataset Description This dataset is a large-scale, high-quality resource designed for training Azerbaijani text embedding models for information retrieval tasks. It contains 671,528 training instances, each consisting of a query, a relevant positive document, and 10 hard-negative documents. The primary goal of this dataset is to facilitate the training of dense retriever models using contrastive learning.… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_retriever_corpus.tabularsentence-similarity100K<n<1M0 likes322 downloads1y agoHugging Face29eddmpython /cleangov-local-budgets 지방재정365 예산 통계·재정지표 Open API 지방재정365 예산 분류의 공식 통계다. 분야·회계·재원별 예산 구성, 교육·인력·의회 관련 경비, 주민 1인당 금액과 사업 비중을 제공한다. 예산 계획을 설명하는 자료이며 실제 지급액이나 결산액으로 해석하지 않는다. 서비스별 제공 항목과 단위는 resources가 소유한다. 재정자립도·재정자주도·통합재정수지비율은 FISCAL-LOC-003의 통합공시에서 제공한다. 출처: https://www.lofin365.go.kr/portal/LF5100000.do 이용 조건: 허브 명세 이용조건 "출처표시, 상업적·비상업적 이용가능, 변형 등 2차적 저작물 작성 가능" (공공누리 1유형 상당) manifest.json은 현재 검증된 파일의 계층·원래 경로·SHA-256·크기·불변 객체 경로를 제공한다. 같은 commit revision으로 manifest와 객체를 내려받고 해시를 검증한다. 예산현액과 지출액은… See the full description on the dataset page: https://huggingface.co/datasets/eddmpython/cleangov-local-budgets.text100K<n<1M0 likes319 downloads16d agoHugging Face30seungheondoh /yt-pe-av-local-embeddingstext10K<n<100K0 likes262 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.