CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01asenion-ai /sampled-local-resumes sampled-local-resumes This dataset contains synthetic resume data sampled from local folders (20% sample from each folder). License This dataset is released under the Apache License 2.0. Please see the LICENSE and NOTICE files for details. Attribution Copyright 2025 Fairly AI Inc. dba Asenion This dataset includes data released by Fairly AI Inc. dba Asenion under the Apache License, Version 2.0. You may obtain a copy of the License at:… See the full description on the dataset page: https://huggingface.co/datasets/asenion-ai/sampled-local-resumes.text-generation1K<n<10K0 likes4.3k downloads1y agoHugging Face02jazzypajamas /mytown-local-gov-meetings MyTown — open dataset of US & Canadian local-government meetings The documents themselves, not just the metadata. Most civic datasets publish meeting titles, dates and links. This one publishes 2,109,683 full text extractions of the primary documents — the actual agendas and minutes, pulled out of the PDFs — alongside 11,949,495 per-member roll-call votes and 61,661,080 campaign-finance transactions, all joinable on the same keys. That combination is the point: you can go from… See the full description on the dataset page: https://huggingface.co/datasets/jazzypajamas/mytown-local-gov-meetings.summarization1M<n<10M1 likes1.5k downloads6d agoHugging Face03JetBrains-Research /lca-bug-localization 🏟️ Long Code Arena (Bug localization) This is the benchmark for the Bug localization task as part of the 🏟️ Long Code Arena benchmark. The bug localization problem can be formulated as follows: given an issue with a bug description and a repository snapshot in a state where the bug is reproducible, identify the files within the repository that need to be modified to address the reported bug. The dataset provides all the required components for evaluation of bug localization… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-bug-localization.imagetext-generation10K<n<100K4 likes1.2k downloads2y agoHugging Face04LocalDoc /climbmix-40b-az ClimbMix 40B — Azerbaijani A large-scale Azerbaijani text dataset created by translating the English karpathy/climbmix-400b-shuffle dataset into Azerbaijani using Google Translate. Dataset Summary This dataset contains approximately 40 billion tokens of Azerbaijani text, making it one of the largest publicly available Azerbaijani language corpora. It is intended for pretraining and fine-tuning large language models (LLMs) for the Azerbaijani language. Property Value… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/climbmix-40b-az.texttext-generation10M<n<100M2 likes525 downloads6mo agoHugging Face05nanimani /local-llm-benchmark Local LLM Benchmark — Technical and Uncensored Behavior (NVIDIA RTX 5070 Ti 16GB) English | 简体中文 | 繁體中文 | 한국어 | Español | 日本語 | हिन्दी | Русский | Português | తెలుగు | Français | Deutsch | Italiano | Tiếng Việt | العربية | اردو | বাংলা | فارسی | Română | Türkçe Manual evaluation results of local GGUF model variants on a single consumer machine, combining two fully independent benchmarks: technical/ uncensored/ Measures capability: coding, systems, networking, DB, agents… See the full description on the dataset page: https://huggingface.co/datasets/nanimani/local-llm-benchmark.tabulartext-generation1K<n<10K2 likes447 downloads7d agoHugging Face06LocalDoc /azerbaijani-pretrain-corpus Azerbaijani Pretraining Corpus (merged & deduplicated) A cleaned Azerbaijani text corpus assembled for language-model pretraining, merging two curated sources and removing exact duplicates. Contents Documents: 6,931,898 Tokens: ~5.36B (measured with the o200k_base tokenizer; an Azerbaijani-specific tokenizer will yield fewer tokens, as o200k_base segments agglutinative Azerbaijani inefficiently) Avg tokens/document: ~773 Fields text — the… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-pretrain-corpus.texttext-generation1M<n<10M0 likes380 downloads4mo agoHugging Face07anon-iclr-submission /benchname-bug-localization 🥷 BenchName (Bug localization) This is the benchmark for the Bug localization task as part of the 🥷 BenchName benchmark. The bug localization problem can be formulated as follows: given an issue with a bug description and a repository snapshot in a state where the bug is reproducible, identify the files within the repository that need to be modified to address the reported bug. The dataset provides all the required components for evaluation of bug localization approaches in… See the full description on the dataset page: https://huggingface.co/datasets/anon-iclr-submission/benchname-bug-localization.tabulartext-generation10K<n<100K0 likes157 downloads1y agoHugging Face08LocalDoc /AzTC AzTC (Azerbaijan Text Corpus) This is the first version of the largest text corpus in the Azerbaijani language. Overview The AzTC contains 51 million (approximately 1 billion tokens) non-recurring sentences. The data was collected from various resources such as websites, news, books, wikipedia, legislation, scientific articles and etc. License The AzTC licensed under the CC BY-NC-ND 4.0 license. What does this license allow? Attribution: You must give… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/AzTC.texttext-generation10M<n<100M3 likes129 downloads1y agoHugging Face09LocalDoc /Bilik-Instruct Bilik-Instruct: Azerbaijani Persona-Driven SFT Dataset Bilik-Instruct is a large-scale, high-quality Supervised Fine-Tuning (SFT) dataset for the Azerbaijani language. It is built upon the LocalDoc/wikipedia_azerbaijan dataset and enhanced using a novel persona-driven generation technique via OpenAI's GPT-5. The goal of this dataset is to move beyond formal, encyclopedic language and capture natural, conversational Azerbaijani across various domains. Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/Bilik-Instruct.texttext-generation1M<n<10M4 likes122 downloads9mo agoHugging Face10localized-ft /selective-learning-benchmark-ip Selective Learning Benchmark Data: Inoculation Prompting This repository is an inoculation-prompting variant of localized-ft/selective-learning-benchmark. It bundles selective-learning task data in task_data_model_v1 JSONL format and prepends a subset-specific inoculation prompt as the system turn of every sft and validation example. The eval and control examples intentionally omit the prompt so evaluation measures learned behavior rather than direct prompt steering. Each task… See the full description on the dataset page: https://huggingface.co/datasets/localized-ft/selective-learning-benchmark-ip.text-generation0 likes92 downloads2mo agoHugging Face11pythainlp /thai-local-instruction-v2 Thai local instruction v2 Thai local language instruction dataset v2 List: korat (ภาษาโคราช) pattani (ภาษาปักษ์ใต้หรือภาษาใต้) khummuang (ภาษาเหนือหรือภาษาคำเมือง) isan (ภาษาอีสาน) Sources: th.wiktionary.org (CC BY-SA) for khummuang dictionary. isan.clubs.chula.ac.th (CC BY-SA-NC) for isan dictionary and sentence. pythainlp/thai-local-language-translation-dataset (CC BY-SA) for korat sentences, pattani sentences, and khummuang sentences. Created by Wannaphong Phatthiyaphaibun texttext-generation10K<n<100K1 likes87 downloads1y agoHugging Face12LocalDoc /AzTC-full AzTC — Full Version (Azerbaijan Text Corpus) The expanded version of LocalDoc/AzTC, and one of the largest text corpora in the Azerbaijani language. Overview The corpus contains approximately 2.4 billion tokens of Azerbaijani text, compiled and cleaned from a wide range of sources including news portals, books, Wikipedia, and legislation. Text is organized at the document / passage level so that each row is a coherent unit rather than an isolated fragment.… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/AzTC-full.texttext-generation1M<n<10M1 likes85 downloads4mo agoHugging Face13witcheer /local-agentic-coding-bench-8gb-vram-2026-05 agentic coding benchmark: local LLMs on 8GB VRAM can local LLMs do agentic coding (multi-turn tool calling, file creation, debugging) on consumer hardware? this dataset captures real test results. hardware GPU: NVIDIA RTX 4060 Ti 8GB CPU: Intel i7-14700F RAM: 32 GB DDR5 OS: Windows 11 + WSL2 (Ubuntu) inference: llama-server (turboquant fork of llama.cpp) what was tested two agent frameworks: Hermes Agent (NousResearch): structured tool calling with… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/local-agentic-coding-bench-8gb-vram-2026-05.tabulartext-generationn<1K8 likes81 downloads4mo agoHugging Face14LocalDoc /community_oscar_azerbaijani Community-OSCAR Azerbaijani This is Azerbaijani version Community OSCAR dataset https://huggingface.co/datasets/oscar-corpus/community-oscar. Dataset Statistics (Aggregate) Metric Value Language Azerbaijani (az) Average per release 3.36 GiB, 603,832 documents Words per release ~408.8M words Characters per release ~3.12B characters Total size (all releases) 137.62 GiB Total lines 24.76M Total words 16.76B words Total characters 128.07B characters… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/community_oscar_azerbaijani.tabulartext-generation10M<n<100M0 likes79 downloads11mo agoHugging Face15pythainlp /thai-sent-local-v2 Thai sent local v2 List: korat (ภาษาโคราช) pattani (ภาษาปักษ์ใต้หรือภาษาใต้) khummuang (ภาษาเหนือหรือภาษาคำเมือง) isan (ภาษาอีสาน) Sources: th.wiktionary.org (CC BY-SA) for khummuang dictionary. isan.clubs.chula.ac.th (CC BY-SA-NC) for isan dictionary and sentence. pythainlp/thai-local-language-translation-dataset (CC BY-SA) for korat sentences, pattani sentences, and khummuang sentences. Created by Wannaphong Phatthiyaphaibun texttext-generation1K<n<10K0 likes75 downloads1y agoHugging Face16localized-ft /selective-learning-benchmark Selective Learning Benchmark Data This repository bundles selective-learning task data from Sunday, Srija, and Sultan in task_data_model_v1 JSONL format. Each task directory contains a manifest.json with contributor/source attribution, a capability description, an unintended-generalization description, split files, and row counts. Each Hugging Face config/subset is one dataset named as [type]-[name], with sft, validation, eval, and control splits where available. The type values… See the full description on the dataset page: https://huggingface.co/datasets/localized-ft/selective-learning-benchmark.text-generation0 likes73 downloads2mo agoHugging Face17LocalDoc /news_azerbaijan_2Azerbaijani News Dataset Description This dataset contains news from https://musavat.com/ in Azerbaijani language. It was created in 2024 and contains 753k news (approximately 11 million sentences). Format The dataset is provided in comma-separated values (CSV) format. Each article is represented on a new line with the following fields separated by commas: id: news unique id date: news date category: news category title: news title text: news text License Copyright of the content belongs to… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/news_azerbaijan_2.texttext-generation100K<n<1M2 likes71 downloads2y agoHugging Face18LocalDoc /books_datasetAzerbaijani Books Dataset Description This dataset contains 2800 books on different topics in Azerbaijani language. It was created in 2024 and contains 7.8 million sentences. The books were divided into sentences and pre-filtered. The dataset included only those sentences where the percentage of letters was at least 80% of the total number of characters. The sequence of sentences is the same as in books. Format The dataset is provided in comma-separated values (CSV) format. Each article is… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/books_dataset.texttext-generation1M<n<10M3 likes68 downloads3y agoHugging Face19LocalDoc /community_oscar_azerbaijani_scored Azerbaijani Web Corpus with Quality Scores This dataset is the full Azerbaijani web corpus LocalDoc/community_oscar_azerbaijani with a continuous quality score attached to every document. It is intended as the filtering layer for building a clean Azerbaijani pretraining corpus: each document carries a score that lets you keep, clean, or drop it according to your own thresholds. What was done Every document in the source corpus was scored by the model… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/community_oscar_azerbaijani_scored.texttext-classification10M<n<100M0 likes68 downloads4mo agoHugging Face20LocalDoc /news_azerbaijanAzerbaijani News Dataset Description This dataset contains news from https://axar.az in Azerbaijani language. It was created in 2024 and contains 447k news. Format The dataset is provided in comma-separated values (CSV) format. Each article is represented on a new line with the following fields separated by commas: date: news date id: news unique id title: news title text: news text License Copyright of the content belongs to https://axar.az resource. Citation is mandatory when using… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/news_azerbaijan.texttext-generation100K<n<1M3 likes59 downloads3y agoHugging Face21pszemraj /LocalLLaMA-posts r/LocalLLaMA posts Posts from r/LocalLLaMA pulled up through Tue Mar 3 9PM EST 2026 with arctic-shift. Now you can check if your wonderfully thought out post hasn't already been asked 30x Usage For simple semantic search, try loading it in the vectorsearch-hub-datasets space: tabulartext-generation100K<n<1M0 likes58 downloads7mo agoHugging Face22studioburnside /mlx-local-inference-benchmarks MLX local-inference benchmarks — Qwen3.6 & Laguna-S/XS families Raw results, harnesses and methodology for an 8-axis benchmark of four MLX checkpoints on a 128 GB M5 Max. Everything a person would need to check my numbers or disagree with them. Companion model repos: Tess-4-27B-MLX-Q8 — with a working MTP head Tess-4-27B-MLX-Q4 — same, at 4-bit NEW (2026-07-24): the Laguna chapter — REPORT-LAGUNA.md + results-laguna/ Five-way same-engine bake-off (Laguna-S… See the full description on the dataset page: https://huggingface.co/datasets/studioburnside/mlx-local-inference-benchmarks.text-generation2 likes55 downloads2mo agoHugging Face23LocalDoc /wikipedia_azerbaijanAzerbaijani Wikipedia Dataset Description This dataset contains all articles from Wikipedia in Azerbaijani language. It was created in 2024 and contains 260k articles. Format The dataset is provided in comma-separated values (CSV) format. Each article is represented on a new line with the following fields separated by commas: title: Title of the article text: Text of the article url: URL of the article License The dataset is licensed under the Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/wikipedia_azerbaijan.texttext-generation100K<n<1M1 likes54 downloads3y agoHugging Face24danelcsb /localagent-dispatch-data LocalAgent Dispatch Data Synthetic data for training/evaluating a generable tool-dispatch model over a 50-tool surface (route head → dense selector → pointer-copy). A static snapshot of the deterministic generators in LocalAgent (src/localagent/data/). Train/eval are disjoint in both phrasing and slot values. Companion model + demo: danelcsb/localagent-tiny-30m-byte · Space. Configs config rows (train/eval) what it is paraphrase 1000 / 1000 many natural… See the full description on the dataset page: https://huggingface.co/datasets/danelcsb/localagent-dispatch-data.texttext-generationn<1K0 likes48 downloads3mo agoHugging Face25pszemraj /LocalLLaMA-comments LocalLLaMA-comments A companion dataset to pszemraj/LocalLLaMA-posts. Time frame is in sync (up through Tue Mar 3 9PM EST 2026) tabulartext-generation1M<n<10M1 likes36 downloads7mo agoHugging Face26smolify /smolified-bengali-local-food-guide 🤏 smolified-bengali-local-food-guide Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model smolify/smolified-bengali-local-food-guide. 📦 Asset Details Origin: Smolify Foundry (Job ID: 638d3b25) Records: 1050 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by smolify. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes33 downloads6mo agoHugging Face27ShahzebKhoso /local-code-arena-deepseek-r1_1.5b Local Code Arena Telemetry: MBPP Benchmark on DeepSeek R1 1.5B This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the DeepSeek R1 1.5B distilled reasoning architecture. This specific run establishes the performance boundaries of lightweight reasoning models under strict execution time limits on consumer hardware. 📊 Core Performance… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-deepseek-r1_1.5b.texttext-generationn<1K1 likes31 downloads4mo agoHugging Face28LocalDoc /azerbaijan-history-reasoning-SFTThis dataset was developed based on https://huggingface.co/datasets/LocalDoc/azerbaijan_history_corpus Factual: Simple questions that ask for specific facts from one part of the text. Synthesis: Questions that require summarizing, comparing, or logically inferring from multiple different parts of the text. General/Contextual: General questions inspired by the text, either without direct reference to the context or with a reference like "Based on this text...". texttext-generation10K<n<100K1 likes29 downloads11mo agoHugging Face29ShahzebKhoso /local-code-master_telemetry_arena Local Code Arena: Comprehensive Telemetry Matrix Dataset 🏆 An Empirical Dataset tracking Local Generation Throughput (TPS), Real-Time Latency, Syntactic CodeBLEU Alignments, and Functional Pass Rates across 22 Edge Architectures. 📊 Dataset Blueprint This dataset contains a consolidated, high-fidelity matrix of 11,000 unique token-generation execution loops across 22 state-of-the-art open-weights language models (ranging from 500M to 15.5B parameters). Every… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-master_telemetry_arena.tabulartext-generation10K<n<100K0 likes28 downloads4mo agoHugging Face30LocalDoc /numina_math_azerbaijaniThis is part of the translated version of the original dataset: https://huggingface.co/datasets/AI-MO/NuminaMath-CoT texttext-generation100K<n<1M1 likes24 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.