datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
isotope-bench
Isotope Bench
An indirect-prompt-injection benchmark for tool-calling agents, plus the
complete audit trail of one recorded run: 438 influence certificates, one for
every action an agent attempted across five defence conditions.
Built for Isotope, which tracks
untrusted influence inside the forward pass. The corpus is independent of that
method and usable with any defence.
💻 Code: https://github.com/NagaYu/isotope
🤗 Demo: https://huggingface.co/spaces/NagaYu/isotope
🤗… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/isotope-bench.deference-keigo-corpus
Deference — Japanese honorific (keigo) error corpus
A corpus for detecting and correcting errors in Japanese honorifics, constructed
mechanically from the norm rather than collected or generated by a model.
The classes, forms and conditions set out in the Council for Cultural Affairs'
report Keigo no Shishin (敬語の指針, 2007) are implemented as rules; correct
sentences are generated from those rules, and documented error types are then
injected — also by rule.
No LLM was… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/deference-keigo-corpus.downstep-bench
Downstep compound pitch-accent benchmark
Japanese noun compounds, their mora segmentation, and every accent their source
dictionaries attest. Built so that "the model has never seen this compound" is a
condition you can actually turn on, rather than a claim you have to trust.
No human annotation is present in this release. Every accent label here is
dictionary-derived. docs/ANNOTATION_GUIDELINES.md ships the protocol, the CSV
format and the agreement statistics for collecting… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/downstep-bench.assay-receipts
Assay Receipt Corpus
Signed internal receipts from an inference provider that is sometimes cheating, together
with the verdict an auditor reached on each one and the ground truth of which model actually
served the request.
Each row is a real receipt, not a summary statistic: it carries the prompt and output token
ids, the JL-projected sketch of the provider's hidden_states, the sign/rank invariants, and
an HMAC signature. With the gpt2 weights you can recompute the sketch… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/assay-receipts.claimcheck-eval
ClaimCheck Evaluation Set
158 hand-checked cases for claim-level groundedness verification: given an
answer and the context that was supplied to the model, is each specific claim in
that answer supported, derived, approximate, unsupported or contradicted?
Bilingual (87 English / 71 Japanese). Every case was run through
ClaimCheck v1.1.0, and the
dataset records what the tool actually returned, not just what it should
have returned.
Cases
158
Languages
English 87 ·… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/claimcheck-eval.Z3-Verified-Reasoning-Graphs
Z3-Verified Constraint Reasoning Dataset
5k Baseline · Production-Ready · Zero Label Noise
The Problem This Solves
Most synthetic reasoning datasets only show the "happy path". Real reasoning requires knowing when to backtrack.
Open-source LLMs hallucinate on constraint satisfaction problems because they are trained on fluent-sounding but logically inconsistent traces. This dataset is different:
❌ No LLM-generated reasoning — zero hallucinations, zero label noise
✅… See the full description on the dataset page: https://huggingface.co/datasets/nagygabor/Z3-Verified-Reasoning-Graphs.scribe-koyobun-usage
Scribe usage-judgment dataset
Span-level data for judging context-dependent kanji/kana usage in Japanese official writing.
Generated by scripts/build_dataset.py in the GitHub repo.
What the claim rests on
The center of this dataset is the hard split: occurrences of words that appear in both usages
(kana and kanji) across the corpus. Because uniform dictionary replacement collapses a word to a single
spelling, it is structurally forced to mislabel one side of this… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/scribe-koyobun-usage.litmus-kernels
Litmus Kernel Verification Corpus
Correct and deliberately-broken Triton kernels, each broken one shipped with
the input that exposes it.
The corpus exists to measure one thing: how much of what a fixed-shape
torch.rand() allclose test calls "correct" actually is. On this corpus the
answer is that 88% of the planted bugs pass that test.
Columns
column
meaning
name
kernel identifier
family
elementwise / reduction / softmax / layernorm / matmul /… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/litmus-kernels.dissent-synthetic-clause-ambiguity
Dissent — synthetic financial clauses with seeded formalization ambiguity
Code · Interactive Space
Every clause in this dataset is synthetic. No real contract, client document or
third-party text was used, quoted or paraphrased. The language follows standard market
forms of drafting so that it is representative of the constructions that cause real
formalization disputes.
What this is for
Autoformalization research is crowded at the producing end and empty at the… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/dissent-synthetic-clause-ambiguity.FactOReS
Description
This dataset contains verifiable claims paired with verification questions and contextual evidence snippets (571 instances in total) in Spanish extracted from online sources.It is designed for the automatic fact-checking task.
The dataset is a novel dataset based on content sourced from Maldita.es, a leading Spanish fact-checking organization.
Each entry includes:
Each entry includes:
claim_id: unique identifier of the claim
claim: textual claim to be verified… See the full description on the dataset page: https://huggingface.co/datasets/nagorebravo/FactOReS.
