datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anle-toaan-gov-vn
Vietnamese Án lệ Corpus — anle.toaan.gov.vn
🇻🇳 Tóm tắt. Bộ dữ liệu các bản án + án lệ Việt Nam thu thập từ cổng
anle.toaan.gov.vn của Tòa án nhân dân tối cao.
Mỗi văn bản đi kèm markdown chuẩn hoá tiếng Việt và một lớp grounding mức câu
(mỗi trích dẫn mang sentence_id + char span trỏ ngược vào markdown). Bộ dữ
liệu là một phần của ViLA common-corpus và ship ba cấu hình HF theo chuẩn
chung: documents (bảng chính) · embeddings (vector 4096-D Nemotron-3-Embed-8B) ·
reduces (toạ… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/anle-toaan-gov-vn.open_government
Open Government Dataset
Open Government is the largest agregation of governement text and data made available as part of open data programs.
In total, the dataset contains approximately 380B tokens. While Open Government aims to become a global resource, in its current state it mostly features open datasets from the US, France, European and international organizations.
The dataset comprises 16 collections curated through two different initiaties: Finance commons and Legal commons.… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/open_government.phapdien-moj-gov-vn
Bộ Pháp Điển Việt Nam — phapdien.moj.gov.vn
🇻🇳 Tóm tắt. Bộ ngữ liệu cấp Điều của Bộ Pháp Điển Việt Nam — bộ pháp điển
chính thức do Bộ Tư pháp công bố. Mỗi dòng documents là một Điều kèm toàn văn đã
chuẩn hoá, chương sở thuộc, đề mục và chủ đề. Kèm theo là vector nhúng ngữ nghĩa 4096-D
(embeddings), toạ độ giảm chiều trong không gian chung ViLA (reduces), và từ điển
ontology song ngữ Việt–Anh (chủ đề · đề mục · thuật ngữ).
🇬🇧 One-line. Article-level corpus of the Bộ Pháp… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/phapdien-moj-gov-vn.gspc-gov
GSPC — governance bank (GovBench)
Bank (governance). Frozen split. Live n is the governance row on GET https://councilof.ai/api/gspc, not a Hub leaderboard score. Not a certificate.
Art 50 dates (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
Council of AI measurement bank. Measurement, not certification.
Live measurement. This bank stands behind the governance row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=governance (family, kind, status and… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-gov.cbba-toaan-gov-vn
Vietnamese Bản án Corpus — congbobanan.toaan.gov.vn
🇻🇳 Tóm tắt. Bản án sơ thẩm/phúc thẩm/giám đốc thẩm/tái thẩm của Việt Nam,
thu thập từ cổng công bố bản án
congbobanan.toaan.gov.vn của Tòa án nhân
dân tối cao. Ba cấu hình HF khoá theo doc_name/id: documents (nội dung
siêu dữ liệu + trích dẫn), embeddings (vector 4096-D Nemotron-3-8B),
reduces (toạ độ t-SNE/UMAP trong không gian chung 6 bộ dữ liệu). Tên
cột và giá trị phân loại bằng tiếng Anh; chỉ nội dung pháp lý giữ tiếng… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/cbba-toaan-gov-vn.gspc-jail-goldbank
GSPC — jail bank (GoldBank-Detector)
Council of AI measurement bank. Measurement, not certification.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
Live measurement. This bank stands behind the jail row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=jail (family, kind, status and n are on that row, never typed here; the… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-jail-goldbank.fattah-golden-superset
Fattah Golden
Fattah Golden is a large-scale, model-agnostic supervised fine-tuning (SFT) superset built by Nomeda Labs to train the Fattah family of coding and agentic coding models.
The dataset is designed as a labeled superset with no baked-in training ratios. This means the stored dataset is the complete cleaned and annotated corpus. Researchers and practitioners choose their own mixture at training time by filtering on the boolean capability columns.
Stats… See the full description on the dataset page: https://huggingface.co/datasets/nomeda-lab/fattah-golden-superset.Goodreads-Books
Dataset Card for "BrightData/Goodreads-Books"
Dataset Summary
Explore a collection of millions of books with the Goodreads dataset, comprising over 6.3M structured records and 14 data fields updated and refreshed regularly.
Each entry includes all major data points such as URLs, book IDs, titles, authors, ratings, number of ratings, reviews, summaries, genres, publication dates, author details and prices.
For a complete list of data points, please refer to the full "Data… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/Goodreads-Books.wildchat-mixed-1k
wildchat-mixed-1k
Real-world chat requests for end-to-end LLM inference benchmarking in fastkernels — Scenario A, the bulk-throughput workload used to saturate continuous batching with a realistic mix of short/long prompts and short/long responses.
What it's for
One dataset that replaces separate prefill-heavy / balanced / decode-heavy splits: its natural length distribution puts prefill-bound and decode-bound requests in the same batch, so a single run yields a… See the full description on the dataset page: https://huggingface.co/datasets/sfc-gh-goliaro/wildchat-mixed-1k.gollem-corpus-2b-pl
GoLLeM Corpus 2B PL
Dokładny korpus treningowy polskiego modelu bazowego
SlayerLab/GoLLeM-110M-PL-v3
(oraz v2) — ten sam zbiór, po którym model przeszedł dwie epoki. Publikujemy go,
aby każdy mógł odtworzyć trening od zera na własnym tokenizerze.
Jak powstał ten plik. Korpus odzyskano przez zdekodowanie stokenizowanego
checkpointu treningowego (byte-level BPE dynaword-32k, round-trip bezstratny; granice
dokumentów = token <|endoftext|>). To jest dokładnie tekst, który model… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/gollem-corpus-2b-pl.goldset
Goldset
Verified bug-fix records for evaluating coding agents.
Every record is a real bug in public software, the fix its author wrote, and the
test that fails before the fix and passes after it. A record is kept only once
both runs have been observed, so what is published is a reproduction rather than
a claim.
896 records from 352 projects, all Python, with
fixes committed between 2010-06-13 and 2026-08-17.
Website · Code and verifier · Datasheet
Quick start
from… See the full description on the dataset page: https://huggingface.co/datasets/goldsetdev/goldset.goodwiki
GoodWiki Dataset
GoodWiki is a 179 million token dataset of English Wikipedia articles collected on September 4, 2023, that have been marked as Good or Featured by Wikipedia editors. The dataset provides these articles in GitHub-flavored Markdown format, preserving layout features like lists, code blocks, math, and block quotes, unlike many other public Wikipedia datasets. Articles are accompanied by a short description of the page as well as any associated categories.
Thanks to a… See the full description on the dataset page: https://huggingface.co/datasets/euirim/goodwiki.longbench-longctx
longbench-longctx
Long-context requests for end-to-end LLM inference benchmarking in fastkernels — Scenario B. Exercises the regimes the bulk set can't reach: long-sequence attention (incl. sparse / sliding-window / DSA), RoPE/YaRN scaling, and large-KV decode.
What it's for
64 real long documents truncated into clean prefill-length buckets from 8K to 128K, each paired with its real multiple-choice question. Prefill-dominated: it measures how kernels scale with… See the full description on the dataset page: https://huggingface.co/datasets/sfc-gh-goliaro/longbench-longctx.samorzad-gov-pl-articles
Artykuły z platformy samorzad.gov.pl
Wersja: v0.2
Zbiór zawiera 81 418 artykułów opublikowanych na stronach 247 instytucji korzystających ze wspólnej platformy samorzad.gov.pl. Są to między innymi urzędy gmin i powiatów, szkoły, instytucje pomocy społecznej i instytucje kultury.
W wersji v0.2 usunięto dane osobowe i kontaktowe z pól tekstowych przeznaczonych dla odbiorcy. Usunięte wartości zastąpiono jednoznacznymi znacznikami, zachowując układ i znaczenie pozostałej treści.… See the full description on the dataset page: https://huggingface.co/datasets/dawidmajewski/samorzad-gov-pl-articles.VeriLoop-Governed-Recurrence-Verified
VLR-Recurrence-Verified
VLR-Recurrence-Verified is a synthetic-data construction release for studying
evidence-convergent program repair. It operationalizes a protected partial order:
a candidate is positive only when it preserves every already-satisfied
obligation and strictly improves at least one unresolved obligation.
Scale
Split
Tasks
Families
Transitions
Balanced pairs
Certified finals
Train
3,500
28
12,250
49,000
3,500
Validation
750
10
2,623… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-Governed-Recurrence-Verified.goodreads-reviews
Goodreads Reviews (deduplicated)
~15,739,967 book reviews scraped from Goodreads, deduplicated.
Columns
Column
Type
Description
user_id
string
Anonymised user hash
book_id
string
Goodreads book ID
review_id
string
Unique review ID
rating
int8
1–5 star rating (0 = no rating)
review_text
string
Full review text
date_added
string
Date added to shelf
date_updated
string
Date last updated
read_at
string
Date finished reading
started_at
string
Date… See the full description on the dataset page: https://huggingface.co/datasets/vngclinh/goodreads-reviews.taskweft-fbd-godot-train
taskweft-fbd-godot-train
Intents and the IEC 61131-3 Function Block Diagrams that carry them out, as an
EditScore-shaped corpus: one root row per intent, three candidates per row (rank1 the
reference diagram, rank3 one that compiles and does the wrong thing, rank5 one the
compiler refuses), and one score row per candidate from the engine itself: api_runner.gd performed the calls on the fixture scene and the returns were read back. Every row is
constructed from a template and a… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/taskweft-fbd-godot-train.Vietverse-SFT-1K-Gold
🇻🇳 Vietverse-SFT (1K Gold Edition)
The "Less is More" Alignment Paradigm for Native Vietnamese Large Language Models
Bộ Dữ Liệu SFT Tiếng Việt Bản Xứ 1.000 Mẫu Gold Tinh Hoa — Chuẩn Mực Căn Chỉnh Mô Hình Ngôn Ngữ
🇻🇳 [Đọc Báo Cáo Kỹ Thuật Tiếng Việt] •
🇬🇧 [Read English Technical Card]
🤗 Hugging Face Dataset • ⚡ Hướng Dẫn Huấn Luyện / Quickstart
🌐 Ngôn Ngữ / Language
📌 Chuyển Hướng Nhanh / Quick Jump… See the full description on the dataset page: https://huggingface.co/datasets/TTP01/Vietverse-SFT-1K-Gold.dsg-state-continuation
DSG State-Continuation
Training data for graph-conditioned long-form fiction generation: given the
narrative state a reader would hold after chapters 1..t-1, and a one-line brief
for chapter t, write chapter t.
Built from 215 public-domain novels (Project Gutenberg, English fiction),
segmented into chapters. 6,876 examples.
Why the state is built this way
The state is not a summary and not a retrieval index. It is a revision-aware
assertion store built causally —… See the full description on the dataset page: https://huggingface.co/datasets/GOVINDFROM/dsg-state-continuation.ukrainian-news-2026
Ukrainian News 2026
Ukrainian-language news articles from 20 national outlets, published between
1 January and 28 August 2026. Extracted body text plus metadata.
Two configs. deduplicated is the default — near-duplicates removed, which
is what you want when mixing this with an already-deduplicated pretraining
corpus. raw is the original release, unchanged.
deduplicated (default)
raw
train-mixin
Documents
419,204
429,427
386,477
Characters
0.97B
1.01B
0.88B
Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Goader/ukrainian-news-2026.ptbr-gov-legal
PT-BR Legal & Government Documents
Part of the MagTina350m pretrain corpus release by Dataseek
under the Magestic.ai brand. This is one of nine silver-layer datasets that fed
dataseek/magtina350m-base.
Summary
935 K Brazilian legal and government documents: federal/state laws, court decisions, regulatory acts, official communications. Mixed corpus combining eduagarcia/LegalPT_dedup (HuggingFace) with a Kaggle Brazilian-legal-proceedings dump.
Source and collection… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-gov-legal.goodreads-books
Goodreads Books Dataset
Dataset Description
A comprehensive dataset of books scraped from Goodreads, including ratings, authors, titles, and various book characteristics.
This dataset contains 3045 books with 20 features each, scraped from Goodreads. It's perfect for:
📚 Book recommendation systems
📊 Literary data analysis
🤖 Machine learning projects
📈 Rating prediction models
🔍 Book discovery algorithms
Dataset Structure
Features… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/goodreads-books.gomodel-go-expert-v4
GoModel Go Expert v4 Dataset
Description
A high-quality dataset for fine-tuning Qwen2.5-Coder-7B to be an expert Go software engineer
with tool-calling capabilities. This is version 4, substantially rebuilt from v3 with:
Structured messages format (not pre-rendered ChatML text)
Go AST-extracted code from real repositories using go/parser
Go 1.26 feature coverage (February 2026 release)
Senior/staff-level engineering content (architecture, distributed systems, API… See the full description on the dataset page: https://huggingface.co/datasets/mencosk/gomodel-go-expert-v4.storyweaver-writing-zh
StoryWeaver 中文写作质量评测集
12 道按写作失效模式反推设计的中文创作题、4 个参赛者写出的 48 篇章节、432 条逐维度两两判决(含裁判完整推理原文)。
来自 StoryWeaver 的写作质量评测轨道。榜单:https://storyweaver.cn/benchmark-writing.html
核心结论
接系统比换一代底模更管用。同一底模接上多 Agent 系统后的胜率:k2.5 **75.1%**、k2.6 **60.2%**;而 k2.5(系统) 对 k2.6(裸) 是 70.3%,反过来只有 37.2%——系统加持能把旧一代底模抬过裸的新一代底模。系统档拿下 22 个维度里的 20 个榜首,包括全部 9 个负向维度。
k2.5 与 k2.6 之间 54.7%,落在噪音带内,不构成结论。
题目怎么设计的
每道题咬住 rubric 里的一个维度或负向维度,用硬约束逼出功力:… See the full description on the dataset page: https://huggingface.co/datasets/godwei123/storyweaver-writing-zh.gohumanize-open-humanizer-dataset
GoHumanize Open Humanizer Dataset
2,957 training pairs and 300 test pairs for teaching a language model to rewrite
AI-styled English prose into natural human writing. Each pair is:
input: a passage rewritten by a large language model in the register typical of LLM output
(formal, smooth, hedged, connective phrases, no contractions);
output: the original human-written passage, from a public-domain book or, since version 2,
from a US federal government publication.
The human… See the full description on the dataset page: https://huggingface.co/datasets/gohumanize/gohumanize-open-humanizer-dataset.ChatGPT-Jailbreak-Prompts-rubend18
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
rfm-rm-as-user-dataset
RFM Reward Model As User Dataset
This dataset was generated for the NeurIPS 2025 paper titled "Capturing Individual Human Preferences with Reward Features". It is released to support the reproducibility of the experiments described in the paper, particularly those in the "Modelling groups of real users" section.
Instead of containing preferences from human raters, this dataset uses 8 publicly available reward models (RMs) as proxies for human raters. This allows for large-scale… See the full description on the dataset page: https://huggingface.co/datasets/google/rfm-rm-as-user-dataset.goldensets
LEGEX Goldensets: Expert-Coded Review-Table Annotations
This repository contains the expert-coded gold annotations for the LEGEX
benchmark of civil-judgment review-table extraction. 1,548 judgments across
19 jurisdictions have been annotated by hand against a shared 14-field schema
covering monetary outcomes, cost allocation, party structure, and industry
classification. Including independent secondary re-annotations, the release
holds 1,974 annotation rows.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/legexbenchmark/goldensets.bureau-docket
The Bureau Corpus — Hyperstition Filings & Records (BIS-F-2026)
This is a work of fiction — a designed art object. Every entry is invented. Nothing here is a prediction, forecast, offer, or advice. The Bureau of Imaginary Solutions does not exist. First filed at ubu.numetal.xyz; companion model at CLINAMEN-45B-A9B.
The public records of the Bureau of Imaginary Solutions — 1,130 documents: 965 hyperstition filings plus 165 Bureau records (memos, appeals, syzygy essays, a… See the full description on the dataset page: https://huggingface.co/datasets/gokhanturhan/bureau-docket.Nigeria_Machinery_Dataset
Nigeria Machinery Usage and Failures Dataset
A structured numeric dataset covering machinery usage rates, equipment failures,
capacity utilization, maintenance costs, and operational downtime across Nigeria's
industrial manufacturing and oil & gas sectors, 2006–2025. It ships
with a companion chain-of-thought reasoning layer derived directly from the
records, for fine-tuning and evaluating LLMs on domain-grounded numeric tasks.
This dataset addresses a real gap: machine-level… See the full description on the dataset page: https://huggingface.co/datasets/gospelgit/Nigeria_Machinery_Dataset.
