CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tmquan /anle-toaan-gov-vn Vietnamese Án lệ Corpus — anle.toaan.gov.vn 🇻🇳 Tóm tắt. Bộ dữ liệu các bản án + án lệ Việt Nam thu thập từ cổng anle.toaan.gov.vn của Tòa án nhân dân tối cao. Mỗi văn bản đi kèm markdown chuẩn hoá tiếng Việt và một lớp grounding mức câu (mỗi trích dẫn mang sentence_id + char span trỏ ngược vào markdown). Bộ dữ liệu là một phần của ViLA common-corpus và ship ba cấu hình HF theo chuẩn chung: documents (bảng chính) · embeddings (vector 4096-D Nemotron-3-Embed-8B) · reduces (toạ… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/anle-toaan-gov-vn.tabulartext-classification10K<n<100K10 likes8.1k downloads8d agoHugging Face02AgentPublic /open_government Open Government Dataset Open Government is the largest agregation of governement text and data made available as part of open data programs. In total, the dataset contains approximately 380B tokens. While Open Government aims to become a global resource, in its current state it mostly features open datasets from the US, France, European and international organizations. The dataset comprises 16 collections curated through two different initiaties: Finance commons and Legal commons.… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/open_government.tabulartext-generation10M<n<100M4 likes2.4k downloads2y agoHugging Face03tmquan /phapdien-moj-gov-vn Bộ Pháp Điển Việt Nam — phapdien.moj.gov.vn 🇻🇳 Tóm tắt. Bộ ngữ liệu cấp Điều của Bộ Pháp Điển Việt Nam — bộ pháp điển chính thức do Bộ Tư pháp công bố. Mỗi dòng documents là một Điều kèm toàn văn đã chuẩn hoá, chương sở thuộc, đề mục và chủ đề. Kèm theo là vector nhúng ngữ nghĩa 4096-D (embeddings), toạ độ giảm chiều trong không gian chung ViLA (reduces), và từ điển ontology song ngữ Việt–Anh (chủ đề · đề mục · thuật ngữ). 🇬🇧 One-line. Article-level corpus of the Bộ Pháp… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/phapdien-moj-gov-vn.imagetext-classification100K<n<1M11 likes1.2k downloads8d agoHugging Face04csoai /gspc-gov GSPC — governance bank (GovBench) Bank (governance). Frozen split. Live n is the governance row on GET https://councilof.ai/api/gspc, not a Hub leaderboard score. Not a certificate. Art 50 dates (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026. Council of AI measurement bank. Measurement, not certification. Live measurement. This bank stands behind the governance row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=governance (family, kind, status and… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-gov.tabularquestion-answeringn<1K0 likes923 downloads2d agoHugging Face05tmquan /cbba-toaan-gov-vn Vietnamese Bản án Corpus — congbobanan.toaan.gov.vn 🇻🇳 Tóm tắt. Bản án sơ thẩm/phúc thẩm/giám đốc thẩm/tái thẩm của Việt Nam, thu thập từ cổng công bố bản án congbobanan.toaan.gov.vn của Tòa án nhân dân tối cao. Ba cấu hình HF khoá theo doc_name/id: documents (nội dung siêu dữ liệu + trích dẫn), embeddings (vector 4096-D Nemotron-3-8B), reduces (toạ độ t-SNE/UMAP trong không gian chung 6 bộ dữ liệu). Tên cột và giá trị phân loại bằng tiếng Anh; chỉ nội dung pháp lý giữ tiếng… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/cbba-toaan-gov-vn.tabulartext-classification1M<n<10M0 likes672 downloads8d agoHugging Face06csoai /gspc-jail-goldbank GSPC — jail bank (GoldBank-Detector) Council of AI measurement bank. Measurement, not certification. Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026. Live measurement. This bank stands behind the jail row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=jail (family, kind, status and n are on that row, never typed here; the… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-jail-goldbank.tabularquestion-answeringn<1K0 likes621 downloads2d agoHugging Face07nomeda-lab /fattah-golden-superset Fattah Golden Fattah Golden is a large-scale, model-agnostic supervised fine-tuning (SFT) superset built by Nomeda Labs to train the Fattah family of coding and agentic coding models. The dataset is designed as a labeled superset with no baked-in training ratios. This means the stored dataset is the complete cleaned and annotated corpus. Researchers and practitioners choose their own mixture at training time by filtering on the boolean capability columns. Stats… See the full description on the dataset page: https://huggingface.co/datasets/nomeda-lab/fattah-golden-superset.tabulartext-generation1M<n<10M1 likes561 downloads4mo agoHugging Face08BrightData /Goodreads-Books Dataset Card for "BrightData/Goodreads-Books" Dataset Summary Explore a collection of millions of books with the Goodreads dataset, comprising over 6.3M structured records and 14 data fields updated and refreshed regularly. Each entry includes all major data points such as URLs, book IDs, titles, authors, ratings, number of ratings, reviews, summaries, genres, publication dates, author details and prices. For a complete list of data points, please refer to the full "Data… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/Goodreads-Books.tabulartext-classification1M<n<10M20 likes395 downloads2y agoHugging Face09sfc-gh-goliaro /wildchat-mixed-1k wildchat-mixed-1k Real-world chat requests for end-to-end LLM inference benchmarking in fastkernels — Scenario A, the bulk-throughput workload used to saturate continuous batching with a realistic mix of short/long prompts and short/long responses. What it's for One dataset that replaces separate prefill-heavy / balanced / decode-heavy splits: its natural length distribution puts prefill-bound and decode-bound requests in the same batch, so a single run yields a… See the full description on the dataset page: https://huggingface.co/datasets/sfc-gh-goliaro/wildchat-mixed-1k.tabulartext-generation1K<n<10K0 likes294 downloads3mo agoHugging Face10SlayerLab /gollem-corpus-2b-pl GoLLeM Corpus 2B PL Dokładny korpus treningowy polskiego modelu bazowego SlayerLab/GoLLeM-110M-PL-v3 (oraz v2) — ten sam zbiór, po którym model przeszedł dwie epoki. Publikujemy go, aby każdy mógł odtworzyć trening od zera na własnym tokenizerze. Jak powstał ten plik. Korpus odzyskano przez zdekodowanie stokenizowanego checkpointu treningowego (byte-level BPE dynaword-32k, round-trip bezstratny; granice dokumentów = token <|endoftext|>). To jest dokładnie tekst, który model… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/gollem-corpus-2b-pl.tabulartext-generation1M<n<10M1 likes277 downloads24d agoHugging Face11goldsetdev /goldset Goldset Verified bug-fix records for evaluating coding agents. Every record is a real bug in public software, the fix its author wrote, and the test that fails before the fix and passes after it. A record is kept only once both runs have been observed, so what is published is a reproduction rather than a claim. 896 records from 352 projects, all Python, with fixes committed between 2010-06-13 and 2026-08-17. Website · Code and verifier · Datasheet Quick start from… See the full description on the dataset page: https://huggingface.co/datasets/goldsetdev/goldset.tabulartext-generationn<1K1 likes271 downloads1mo agoHugging Face12euirim /goodwiki GoodWiki Dataset GoodWiki is a 179 million token dataset of English Wikipedia articles collected on September 4, 2023, that have been marked as Good or Featured by Wikipedia editors. The dataset provides these articles in GitHub-flavored Markdown format, preserving layout features like lists, code blocks, math, and block quotes, unlike many other public Wikipedia datasets. Articles are accompanied by a short description of the page as well as any associated categories. Thanks to a… See the full description on the dataset page: https://huggingface.co/datasets/euirim/goodwiki.tabulartext-generation10K<n<100K55 likes217 downloads3y agoHugging Face13sfc-gh-goliaro /longbench-longctx longbench-longctx Long-context requests for end-to-end LLM inference benchmarking in fastkernels — Scenario B. Exercises the regimes the bulk set can't reach: long-sequence attention (incl. sparse / sliding-window / DSA), RoPE/YaRN scaling, and large-KV decode. What it's for 64 real long documents truncated into clean prefill-length buckets from 8K to 128K, each paired with its real multiple-choice question. Prefill-dominated: it measures how kernels scale with… See the full description on the dataset page: https://huggingface.co/datasets/sfc-gh-goliaro/longbench-longctx.tabulartext-generationn<1K0 likes198 downloads3mo agoHugging Face14dawidmajewski /samorzad-gov-pl-articles Artykuły z platformy samorzad.gov.pl Wersja: v0.2 Zbiór zawiera 81 418 artykułów opublikowanych na stronach 247 instytucji korzystających ze wspólnej platformy samorzad.gov.pl. Są to między innymi urzędy gmin i powiatów, szkoły, instytucje pomocy społecznej i instytucje kultury. W wersji v0.2 usunięto dane osobowe i kontaktowe z pól tekstowych przeznaczonych dla odbiorcy. Usunięte wartości zastąpiono jednoznacznymi znacznikami, zachowując układ i znaczenie pozostałej treści.… See the full description on the dataset page: https://huggingface.co/datasets/dawidmajewski/samorzad-gov-pl-articles.tabulartext-generation10K<n<100K0 likes168 downloads1mo agoHugging Face15tsinghua-sigs-robot-lab /VeriLoop-Governed-Recurrence-Verified VLR-Recurrence-Verified VLR-Recurrence-Verified is a synthetic-data construction release for studying evidence-convergent program repair. It operationalizes a protected partial order: a candidate is positive only when it preserves every already-satisfied obligation and strictly improves at least one unresolved obligation. Scale Split Tasks Families Transitions Balanced pairs Certified finals Train 3,500 28 12,250 49,000 3,500 Validation 750 10 2,623… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-Governed-Recurrence-Verified.tabulartext-generation100K<n<1M1 likes165 downloads1mo agoHugging Face16vngclinh /goodreads-reviews Goodreads Reviews (deduplicated) ~15,739,967 book reviews scraped from Goodreads, deduplicated. Columns Column Type Description user_id string Anonymised user hash book_id string Goodreads book ID review_id string Unique review ID rating int8 1–5 star rating (0 = no rating) review_text string Full review text date_added string Date added to shelf date_updated string Date last updated read_at string Date finished reading started_at string Date… See the full description on the dataset page: https://huggingface.co/datasets/vngclinh/goodreads-reviews.tabulartext-classification10M<n<100M0 likes148 downloads5mo agoHugging Face17chibifire /taskweft-fbd-godot-train taskweft-fbd-godot-train Intents and the IEC 61131-3 Function Block Diagrams that carry them out, as an EditScore-shaped corpus: one root row per intent, three candidates per row (rank1 the reference diagram, rank3 one that compiles and does the wrong thing, rank5 one the compiler refuses), and one score row per candidate from the engine itself: api_runner.gd performed the calls on the fixture scene and the returns were read back. Every row is constructed from a template and a… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/taskweft-fbd-godot-train.tabulartext-generation10K<n<100K0 likes113 downloads16d agoHugging Face18TTP01 /Vietverse-SFT-1K-Gold 🇻🇳 Vietverse-SFT (1K Gold Edition) The "Less is More" Alignment Paradigm for Native Vietnamese Large Language Models Bộ Dữ Liệu SFT Tiếng Việt Bản Xứ 1.000 Mẫu Gold Tinh Hoa — Chuẩn Mực Căn Chỉnh Mô Hình Ngôn Ngữ 🇻🇳 [Đọc Báo Cáo Kỹ Thuật Tiếng Việt] &nbsp;•&nbsp; 🇬🇧 [Read English Technical Card] 🤗 Hugging Face Dataset • ⚡ Hướng Dẫn Huấn Luyện / Quickstart 🌐 Ngôn Ngữ / Language 📌 Chuyển Hướng Nhanh / Quick Jump… See the full description on the dataset page: https://huggingface.co/datasets/TTP01/Vietverse-SFT-1K-Gold.tabulartext-generation1K<n<10K4 likes110 downloads1mo agoHugging Face19GOVINDFROM /dsg-state-continuation DSG State-Continuation Training data for graph-conditioned long-form fiction generation: given the narrative state a reader would hold after chapters 1..t-1, and a one-line brief for chapter t, write chapter t. Built from 215 public-domain novels (Project Gutenberg, English fiction), segmented into chapters. 6,876 examples. Why the state is built this way The state is not a summary and not a retrieval index. It is a revision-aware assertion store built causally —… See the full description on the dataset page: https://huggingface.co/datasets/GOVINDFROM/dsg-state-continuation.tabulartext-generation1K<n<10K0 likes103 downloads20d agoHugging Face20Goader /ukrainian-news-2026 Ukrainian News 2026 Ukrainian-language news articles from 20 national outlets, published between 1 January and 28 August 2026. Extracted body text plus metadata. Two configs. deduplicated is the default — near-duplicates removed, which is what you want when mixing this with an already-deduplicated pretraining corpus. raw is the original release, unchanged. deduplicated (default) raw train-mixin Documents 419,204 429,427 386,477 Characters 0.97B 1.01B 0.88B Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Goader/ukrainian-news-2026.imagetext-generation1M<n<10M0 likes93 downloads25d agoHugging Face21dataseek /ptbr-gov-legal PT-BR Legal & Government Documents Part of the MagTina350m pretrain corpus release by Dataseek under the Magestic.ai brand. This is one of nine silver-layer datasets that fed dataseek/magtina350m-base. Summary 935 K Brazilian legal and government documents: federal/state laws, court decisions, regulatory acts, official communications. Mixed corpus combining eduagarcia/LegalPT_dedup (HuggingFace) with a Kaggle Brazilian-legal-proceedings dump. Source and collection… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-gov-legal.tabulartext-generation100K<n<1M0 likes87 downloads5mo agoHugging Face22codealchemist01 /goodreads-books Goodreads Books Dataset Dataset Description A comprehensive dataset of books scraped from Goodreads, including ratings, authors, titles, and various book characteristics. This dataset contains 3045 books with 20 features each, scraped from Goodreads. It's perfect for: 📚 Book recommendation systems 📊 Literary data analysis 🤖 Machine learning projects 📈 Rating prediction models 🔍 Book discovery algorithms Dataset Structure Features… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/goodreads-books.tabulartext-classification1K<n<10K0 likes82 downloads11mo agoHugging Face23mencosk /gomodel-go-expert-v4 GoModel Go Expert v4 Dataset Description A high-quality dataset for fine-tuning Qwen2.5-Coder-7B to be an expert Go software engineer with tool-calling capabilities. This is version 4, substantially rebuilt from v3 with: Structured messages format (not pre-rendered ChatML text) Go AST-extracted code from real repositories using go/parser Go 1.26 feature coverage (February 2026 release) Senior/staff-level engineering content (architecture, distributed systems, API… See the full description on the dataset page: https://huggingface.co/datasets/mencosk/gomodel-go-expert-v4.tabulartext-generation10K<n<100K0 likes76 downloads2mo agoHugging Face24godwei123 /storyweaver-writing-zh StoryWeaver 中文写作质量评测集 12 道按写作失效模式反推设计的中文创作题、4 个参赛者写出的 48 篇章节、432 条逐维度两两判决(含裁判完整推理原文)。 来自 StoryWeaver 的写作质量评测轨道。榜单:https://storyweaver.cn/benchmark-writing.html 核心结论 接系统比换一代底模更管用。同一底模接上多 Agent 系统后的胜率:k2.5 **75.1%**、k2.6 **60.2%**;而 k2.5(系统) 对 k2.6(裸) 是 70.3%,反过来只有 37.2%——系统加持能把旧一代底模抬过裸的新一代底模。系统档拿下 22 个维度里的 20 个榜首,包括全部 9 个负向维度。 k2.5 与 k2.6 之间 54.7%,落在噪音带内,不构成结论。 题目怎么设计的 每道题咬住 rubric 里的一个维度或负向维度,用硬约束逼出功力:… See the full description on the dataset page: https://huggingface.co/datasets/godwei123/storyweaver-writing-zh.tabulartext-generationn<1K1 likes75 downloads2mo agoHugging Face25gohumanize /gohumanize-open-humanizer-dataset GoHumanize Open Humanizer Dataset 2,957 training pairs and 300 test pairs for teaching a language model to rewrite AI-styled English prose into natural human writing. Each pair is: input: a passage rewritten by a large language model in the register typical of LLM output (formal, smooth, hedged, connective phrases, no contractions); output: the original human-written passage, from a public-domain book or, since version 2, from a US federal government publication. The human… See the full description on the dataset page: https://huggingface.co/datasets/gohumanize/gohumanize-open-humanizer-dataset.tabulartext-generation1K<n<10K0 likes74 downloads1d agoHugging Face26philosopher-from-god /ChatGPT-Jailbreak-Prompts-rubend18 Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K2 likes73 downloads1y agoHugging Face27google /rfm-rm-as-user-dataset RFM Reward Model As User Dataset This dataset was generated for the NeurIPS 2025 paper titled "Capturing Individual Human Preferences with Reward Features". It is released to support the reproducibility of the experiments described in the paper, particularly those in the "Modelling groups of real users" section. Instead of containing preferences from human raters, this dataset uses 8 publicly available reward models (RMs) as proxies for human raters. This allows for large-scale… See the full description on the dataset page: https://huggingface.co/datasets/google/rfm-rm-as-user-dataset.tabulartext-generation10K<n<100K10 likes70 downloads11mo agoHugging Face28legexbenchmark /goldensets LEGEX Goldensets: Expert-Coded Review-Table Annotations This repository contains the expert-coded gold annotations for the LEGEX benchmark of civil-judgment review-table extraction. 1,548 judgments across 19 jurisdictions have been annotated by hand against a shared 14-field schema covering monetary outcomes, cost allocation, party structure, and industry classification. Including independent secondary re-annotations, the release holds 1,974 annotation rows. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/legexbenchmark/goldensets.tabulartext-classification1K<n<10K0 likes68 downloads2mo agoHugging Face29gokhanturhan /bureau-docket The Bureau Corpus — Hyperstition Filings & Records (BIS-F-2026) This is a work of fiction — a designed art object. Every entry is invented. Nothing here is a prediction, forecast, offer, or advice. The Bureau of Imaginary Solutions does not exist. First filed at ubu.numetal.xyz; companion model at CLINAMEN-45B-A9B. The public records of the Bureau of Imaginary Solutions — 1,130 documents: 965 hyperstition filings plus 165 Bureau records (memos, appeals, syzygy essays, a… See the full description on the dataset page: https://huggingface.co/datasets/gokhanturhan/bureau-docket.tabulartext-generation1K<n<10K0 likes62 downloads2mo agoHugging Face30gospelgit /Nigeria_Machinery_Dataset Nigeria Machinery Usage and Failures Dataset A structured numeric dataset covering machinery usage rates, equipment failures, capacity utilization, maintenance costs, and operational downtime across Nigeria's industrial manufacturing and oil & gas sectors, 2006–2025. It ships with a companion chain-of-thought reasoning layer derived directly from the records, for fine-tuning and evaluating LLMs on domain-grounded numeric tasks. This dataset addresses a real gap: machine-level… See the full description on the dataset page: https://huggingface.co/datasets/gospelgit/Nigeria_Machinery_Dataset.tabulartabular-classificationn<1K0 likes60 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.