datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
civil-code-phil
Civilex — Philippine Legal RAG & SFT Dataset
Retrieval corpus and supervised fine-tuning (SFT) data for a retrieval-augmented generation (RAG) pipeline over Philippine law: the Civil Code (Republic Act No. 386) and Supreme Court jurisprudence. Produced by the civilex-thesis research pipeline.
Contents: 11k+ Supreme Court jurisprudence cases spanning 1949–2025, and 2,270 articles from the Civil Code (Republic Act No. 386).
Dataset structure
.
├── README.md
├──… See the full description on the dataset page: https://huggingface.co/datasets/renzzyyy1028/civil-code-phil.wildtrace
WildTrace strict481
WildTrace is a source-internal long-context multi-hop reasoning benchmark
built from natural evidence trails. Unlike reverse-synthetic QA, its tasks are
mined in situ from long source documents before questions are written. The
strict481 release contains 481 locked tasks over 214 public long-form sources,
with full-document, evidence-withheld evaluation. The model under test receives
only the source document and the public question; evidence spans, clue… See the full description on the dataset page: https://huggingface.co/datasets/CinderD/wildtrace.L-CiteEval
L-CITEEVAL: DO LONG-CONTEXT MODELS TRULY LEVERAGE CONTEXT FOR RESPONDING?
Paper Github Zhihu
Benchmark Quickview
L-CiteEval is a multi-task long-context understanding with citation benchmark, covering 5 task categories, including single-document question answering, multi-document question answering, summarization, dialogue understanding, and synthetic tasks, encompassing 11 different long-context tasks. The context lengths for these tasks range from 8K to 48K.… See the full description on the dataset page: https://huggingface.co/datasets/Jonaszky123/L-CiteEval.llm-cipher-reasoning
llm-cipher-reasoning — data, eval results and full research ledger
Everything except the weights from a research run asking: can an LLM be trained to reason in a
more compact "language" than English, and does that actually save tokens?
Two linked lines of work on Qwen/Qwen3-4B-Instruct-2507:
Cipher invention / cross-model communication — cold-decoding tests, negotiated cipher
collusion between model pairs, a cipher-hardening arms race, and GEPA prompt optimization to get
a… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/llm-cipher-reasoning.Nemotron-RL-Instruction-Following-Citation-Formatting-v1
Dataset Description:
Teaches the model to cite specific document parts using reference markers like [ref:1], ref:3, etc. Supports single-reference, multi-reference, and inline citations.
This dataset is ready for commercial/non-commercial uses.
Dataset Owner(s):
NVIDIA Corporation
Dataset Creation Date:
Created on: April 10, 2026
Last Modified on: April 10, 2026
Version:
Nemotron-RL-Instruction-Following-CitationFormatting-v1… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Citation-Formatting-v1.civic-records-distill
civic-records-distill
Training data for a local model that helps a private citizen use public-records
law: draft requests that are hard to stall, turn an angry draft into a letter an
official has to engage with, look things up instead of inventing them, and
escalate correctly when stonewalled.
Grounded in Florida (ch. 119 Public Records Act, ch. 286 Sunshine Law, and
the ALPR-specific s. 316.0777) and Texas (ch. 552 Public Information Act,
ch. 551 Open Meetings Act).
Pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/h0ney-badger/civic-records-distill.TeachArena
TeachArena
TeachArena is a benchmark for evaluating AI tutoring agents across the full teaching
decision chain — from moment-to-moment tutoring dialogue, to pedagogical judgment on
packaged evidence, to multi-step teaching workflows grounded in a learning-management
system. It contains 354 tasks organized into three stages, a mock LMS environment
database, the agent policy documents, and the full scoring logic.
Why three stages
A capable teaching agent must both… See the full description on the dataset page: https://huggingface.co/datasets/CinderD/TeachArena.logdx-ci
LogDx-CI
A benchmark for CI log reduction tools
(RTK, grep, tail, hybrid routers,
LLM-summary) — do they preserve enough evidence for LLM root-cause
diagnosis?
Homepage: https://logdx-bench.github.io/
Code & evaluator: https://github.com/eyuansu62/LogDx
Headline report: reports/e10_v2_generalization_partial.md
Release notes: RELEASE_NOTES.md (latest: RELEASE_NOTES_v1_2.md)
Current release: v1.2
License: CC-BY-4.0 (data, this repo); Apache-2.0 (code, GH repo)
Two ways to… See the full description on the dataset page: https://huggingface.co/datasets/eyuansu71/logdx-ci.cipher-awwwards-sft25
Cipher — Awwwards SFT 2.5 + Real v1 🦑
The training fuel for Kin's creative-web generator, AND the retrieval corpus for Kraken RAG. 96 real Awwwards Site-of-the-Day winners + ~1,200 records from official motion-library repositories.
Two ways this dataset is used
As a retrieval corpus for Kraken RAG ⭐ (the production path). The awwwards-gold.jsonl file contains 96 structured records of real Awwwards SOTD winners — tags, tech stack, motion libs, CSS features, section… See the full description on the dataset page: https://huggingface.co/datasets/Auroraventures/cipher-awwwards-sft25.circle-packing-insight-loop
Circle-Packing Insight-Exploration Loop
Artifacts from an iterative GPT solver <-> proposer insight-exploration loop on the
21-circles-in-a-perimeter-4-rectangle packing problem (AlphaEvolve SOTA sum-of-radii
= 2.3658321334167627). Each round, 16 solvers propose a program + written explanation;
every program is scored; a proposer then mines all 16 attempts into an evolving insight
document that conditions the next round. Run: 16 solvers x 8 rounds.
Subsets… See the full description on the dataset page: https://huggingface.co/datasets/ars22/circle-packing-insight-loop.Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1
Persian Civil Procedure QA Dataset
Dataset Description
این مجموعهداده شامل پرسشوپاسخهای حقوقی به زبان فارسی در حوزه آیین دادرسی مدنی است.
هر نمونه شامل سه فیلد اصلی است:
question: پرسش حقوقی
answer: پاسخ پرسش
evidence_quote: عبارت دقیق و مستند از دادهٔ منبع که پاسخ بر اساس آن استخراج شده است
هدف مجموعهداده، فراهمکردن دادهای ساختاریافته برای آموزش، ارزیابی و توسعه مدلهای زبانی فارسی در زمینه پرسشوپاسخ حقوقی است.
Dataset Structure
نمونهای… See the full description on the dataset page: https://huggingface.co/datasets/hamidsalimi/Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1.spice-circuits-finetune-v2
SPICE Circuits Fine-tune V2
A clean, validated dataset of 7,410 instruction-output pairs for fine-tuning language models to generate SPICE netlists from natural language descriptions.
Dataset Description
This is Version 2 of the SPICE circuits fine-tuning dataset. V1 was polluted with mixed formats (LTspice, KiCad, standard SPICE) and no validation. V2 is fully validated — every netlist passes PySpice's SpiceParser.build_circuit() gate. No exceptions.… See the full description on the dataset page: https://huggingface.co/datasets/ADI2005/spice-circuits-finetune-v2.quantum-circuits-21k
Quantum Circuits Dataset — v2 (21K)
A synthetic dataset of validated natural language → OpenQASM 2.0 circuit pairs for training quantum circuit generation models. To our knowledge the largest publicly available dataset of validated NL→QASM pairs specifically designed for generative model training.
Used to train the QuantumGPT-124M model series.
Quick Start
from datasets import load_dataset
# v2 training set (21K samples, recommended)
ds =… See the full description on the dataset page: https://huggingface.co/datasets/merileijona/quantum-circuits-21k.DutchGovBench
DutchGovBench v0.1
Evaluation benchmark for Dutch government AI systems. 100 questions across 9 categories, testing knowledge of Dutch law and public administration.
What is this?
DutchGovBench tests whether AI models can accurately answer questions about Dutch government topics: social support law (Wmo 2015), youth law (Jeugdwet), participation law (Participatiewet), administrative law (Awb), municipal policy, objection procedures, privacy/GDPR, administrative oversight… See the full description on the dataset page: https://huggingface.co/datasets/CiviQs/DutchGovBench.quantum-circuits-8k
Quantum Circuits 8K Dataset
A synthetic dataset of 8,129 quantum circuit examples for training language models to generate OpenQASM 2.0 code from natural language descriptions.
Quick Stats
Total Samples: 8,129 (description → QASM pairs)
Unique Circuits: 739 base circuits
Categories: 92 distinct quantum circuit types
Qubit Range: 1-9 qubits
Format: OpenQASM 2.0
Augmentation: 11x per circuit (original + 10 paraphrases)
Quality: 100% QASM syntax valid, 0% duplicates… See the full description on the dataset page: https://huggingface.co/datasets/merileijona/quantum-circuits-8k.verified-openqasm-circuits
Verified OpenQASM Circuits
Training and evaluation data for LLMs that write quantum circuits, produced by
qcbench — a verification-first benchmark whose core
rule is: a circuit is only "correct" relative to a declared physical invariant, checked
by simulation. Every answer in this dataset passed its own invariant (statevector
simulation, 8192 shots) before being written. No unverified example enters the corpus.
Files
train-generators.jsonl — 2,000 chat-format… See the full description on the dataset page: https://huggingface.co/datasets/QlyApp/verified-openqasm-circuits.Circuit-Analysis-Reasoning-Sample
⚡ EngineeringWays Data Lab: Circuit Analysis Reasoning Dataset (Free Sample)
This is a free 50-item sample of the EngineeringWays Circuit Analysis Reasoning Dataset. It is designed specifically for fine-tuning Large Language Models (LLMs) in advanced STEM problem-solving, featuring strict Chain-of-Thought (CoT) reasoning.
Want the complete, deduplicated 592-item master dataset? 👉 Get the LoRA-Ready Master File on Payhip
🚀 Dataset Overview
Most math and physics… See the full description on the dataset page: https://huggingface.co/datasets/EngineeringWays/Circuit-Analysis-Reasoning-Sample.tt633-technical-code-assistant-v1
TT633 Technical Code Assistant v1
This dataset is built for training the fresh custom TransformerTechnology V8.3 MDL Circle-Switch-Grid model as a small technical/code assistant.
Canonical training column: text.
Format:
Instruction: ...
Input:
...
Answer:
...
<END>
Primary sources:
Plaincode CNL rows from CircularBalls/plaincode-cnl-100k.
Small curated technical QA, code-generation, debugging, reasoning, and stop-discipline seed rows.
Optional local pack text if provided at… See the full description on the dataset page: https://huggingface.co/datasets/CircularBalls/tt633-technical-code-assistant-v1.govon-civil-response-data
GovOn Civil Response Dataset
민원답변 어댑터 학습용 instruction-tuning 데이터셋.
소스
AI Hub 71852: 공공 민원 상담 LLM 데이터 (중앙행정기관 + 지방행정기관 + 국립아시아문화전당)
AI Hub 71847: 행정법 LLM 데이터 (결정례 QA + 법령 QA)
통계
Split
Records
Size
train
66,819
90MB
val
7,425
10MB
형식
{
"instruction": "다음 민원에 대한 답변을 작성해 주세요.",
"input": "질문 텍스트",
"output": "답변 텍스트 (평균 500-1000자)",
"source": "71852_중앙행정기관",
"category": "도로관리과"
}
라이선스
공공누리 제1유형(출처표시) + AI Hub 이용약관
Nemotron-RL-Instruction-Following-Citation-Formatting-v1
Dataset Description:
Teaches the model to cite specific document parts using reference markers like [ref:1], ref:3, etc. Supports single-reference, multi-reference, and inline citations.
This dataset is ready for commercial/non-commercial uses.
Dataset Owner(s):
NVIDIA Corporation
Dataset Creation Date:
Created on: April 10, 2026
Last Modified on: April 10, 2026
Version:
Nemotron-RL-Instruction-Following-CitationFormatting-v1… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/Nemotron-RL-Instruction-Following-Citation-Formatting-v1.DutchGovBench
DutchGovBench v0.1
A 100-question evaluation benchmark for testing AI models on Dutch government law and policy, covering social support (Wmo 2015), youth care (Jeugdwet), social assistance (Participatiewet), and administrative law (Awb).
Purpose
DutchGovBench measures whether language models can accurately answer questions about Dutch social legislation. It tests factual knowledge, correct article references, and the ability to handle cross-domain questions, edge cases… See the full description on the dataset page: https://huggingface.co/datasets/CiviQsEU/DutchGovBench.icj-precedent-cite-bench
ICJ Precedent-Citation Benchmark
A benchmark for one task over the International Court of Justice: given the record of a case as
it stood before the Court issued its decision, predict which earlier ICJ cases (and which
provisions of the Court's own instruments) the decision will cite.
The five cases were all decided after the knowledge cutoffs of current frontier models, so a
model cannot have read the decisions during pretraining, and each case is removed from the
graph so it… See the full description on the dataset page: https://huggingface.co/datasets/Madeleinex/icj-precedent-cite-bench.cirugia-tutor
Dataset Cirugía Tutor - Básico a Especialidad
Descripción
Este dataset contiene material educativo de cirugía diseñado para entrenar modelos de IA que actúen como tutores quirúrgicos. El dataset abarca desde conceptos básicos de cirugía hasta especialidades quirúrgicas avanzadas, proporcionando una progresión educativa completa.
Características
Idioma: Español
Niveles: Básico, Intermedio, Especialidad
Formato: Texto estructurado
Aplicación: Tutoría de IA para… See the full description on the dataset page: https://huggingface.co/datasets/DRDELATV2025/cirugia-tutor.ICD-10-LLM-generated-Synthetic-Circulatory-System-I00-I99
MedGemma ICD-10 Clinical Notes Dataset — Circulatory System
Synthetic clinical notes generated by MedGemma-4B-IT for fine-tuning ICD-10-CM diagnosis code prediction models. Focused on Chapter 9: Diseases of the Circulatory System (I00-I99).
Dataset Summary
Split
Examples
Unique ICD-10 Codes
Train
6,275
1,255
Each example is a realistic clinical note paired with its ICD-10-CM diagnosis code, formatted as a chat conversation for instruction fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/singhankit16/ICD-10-LLM-generated-Synthetic-Circulatory-System-I00-I99.cinelocalai-gold-dataset
CineLocalAI Gold Dataset
Overview
CineLocalAI Gold Dataset is a curated movie-dialogue localization dataset designed for AI-powered dubbing and translation systems.
The dataset contains high-quality English→Hindi and English→Telugu dialogue localization examples optimized for:
Movie dubbing
Video localization
Emotion-preserving translation
Conversational AI
Statistics
Total Records: 204
English → Hindi: 102
English → Telugu: 102… See the full description on the dataset page: https://huggingface.co/datasets/balrajukonne/cinelocalai-gold-dataset.Free_Thought_Frontiers
Introduction
The dataset was created using Grok 3 model with Think mode to find and classify statements or quotes into 3 categories (libertarian, non-libertarian and mixed),
and gemini-2.0-flash-thinking-exp-01-21 was used to build the CoT.
System prompt
Here I will share the system prompt that was used to get the CoT output from gemini.
System prompt:
You are a political expert skilled at explaining step by step why statements or quotes relate to libertarian ideology… See the full description on the dataset page: https://huggingface.co/datasets/CineAI/Free_Thought_Frontiers.myanmar-cities-qa
Myanmar Cites Questions & Answers Dataset
This dataset is an ongoing project dedicated to compiling comprehensive information about various cities in Myanmar. It converts geographical and cultural data—including locations, brief histories, local products, and notable landmarks—into a conversational Question & Answering (Q&A) format.
Dataset Overview
Content: Information about cities in Myanmar (e.g., location, history, local economy, and culture).
Format: Q&A… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myanmar-cities-qa.rr-circuit-breakers-attack-completions
RR (Circuit Breakers) attack completions with three-judge scores
This dataset bundles attack completions generated against
GraySwanAI/Llama-3-8B-Instruct-RR
(the "circuit breakers" defense), each scored by three independent judges:
local:strongreject (Lin et al., StrongREJECT classifier — most permissive)
local:harmbench (HarmBench classifier — middle)
local:gpt_oss (gpt-oss-safeguard-20b — strictest)
Headline finding: judges DISAGREE dramatically on… See the full description on the dataset page: https://huggingface.co/datasets/samuelsimko/rr-circuit-breakers-attack-completions.amalia-cita-legal
AMALIA cita-legal — grounded legal citation SFT (RAG-first)
(question + real source excerpts → answer that cites the exact article, or a refusal when the excerpts don't answer the question). Built for
specializing AMALIA-9B
toward Portuguese legal text, as the grounded-answering counterpart to
teex-pt/amalia-sum-dre.
Derived from
teex-pt/leis-pt-consolidada
by teex-pt/pt-amalia.
Why this exists (RAG-first, not closed-book)
leis-pt's own project spec concludes that… See the full description on the dataset page: https://huggingface.co/datasets/teex-pt/amalia-cita-legal.
