datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
latentsSO-101_dataset03_20260827_145444This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/nagaenaga/SO-101_dataset03_20260827_145444.so-101_dataset02_20260827_130357This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/nagaenaga/so-101_dataset02_20260827_130357.saccade-egomotion-bench
Saccade ego-motion benchmark
The stream, the raw decision signals, and the per-frame measurements behind
Saccade — an always-on edge VLM that re-encodes only
the image patches whose change ego-motion cannot explain.
This dataset exists so the central claim can be checked without running our code.
💻 Code: https://github.com/NagaYu/saccade
🤖 Model: https://huggingface.co/NagaYu/saccade-predictor
🚀 Demo: https://huggingface.co/spaces/NagaYu/saccade
The claim, in… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/saccade-egomotion-bench.isotope-bench
Isotope Bench
An indirect-prompt-injection benchmark for tool-calling agents, plus the
complete audit trail of one recorded run: 438 influence certificates, one for
every action an agent attempted across five defence conditions.
Built for Isotope, which tracks
untrusted influence inside the forward pass. The corpus is independent of that
method and usable with any defence.
💻 Code: https://github.com/NagaYu/isotope
🤗 Demo: https://huggingface.co/spaces/NagaYu/isotope
🤗… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/isotope-bench.bleep-spans
Bleep spans — synthetic sensitive-speech regions with frame-accurate labels
Where sensitive information is spoken, and what kind it is — never what was
said.
Every recording is synthetic. No real telephone call, clinical recording, or any
other real speech was used, recorded, or derived from at any stage.
🤗 Model: NagaYu/bleep-0.09b
🎛️ Demo: NagaYu/bleep
What a row contains
utt_id, voice_key, condition, duration, subsets, and three parallel
arrays —… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/bleep-spans.deference-keigo-corpus
Deference — Japanese honorific (keigo) error corpus
A corpus for detecting and correcting errors in Japanese honorifics, constructed
mechanically from the norm rather than collected or generated by a model.
The classes, forms and conditions set out in the Council for Cultural Affairs'
report Keigo no Shishin (敬語の指針, 2007) are implemented as rules; correct
sentences are generated from those rules, and documented error types are then
injected — also by rule.
No LLM was… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/deference-keigo-corpus.downstep-bench
Downstep compound pitch-accent benchmark
Japanese noun compounds, their mora segmentation, and every accent their source
dictionaries attest. Built so that "the model has never seen this compound" is a
condition you can actually turn on, rather than a claim you have to trust.
No human annotation is present in this release. Every accent label here is
dictionary-derived. docs/ANNOTATION_GUIDELINES.md ships the protocol, the CSV
format and the agreement statistics for collecting… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/downstep-bench.rebar-structure
🧱 Rebar Structure
A corpus for restoring the heading hierarchy of Japanese documents from flat text and
for evaluating structure-aware chunking. Each record is a flattened document, its
gold heading tree (positions + depths), and a damaged variant simulating PDF/text
extraction.
Code: https://github.com/NagaYu/rebar
Model: https://huggingface.co/NagaYu/rebar-heading-classifier
Demo (Space): https://huggingface.co/spaces/NagaYu/rebar
Why it exists
The same… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/rebar-structure.kvine-agent-trajectory-bench
KVine agent-trajectory benchmark & resident-set policy training data 🌿
Two things from the KVine research prototype
(branch-aware KV cache for agent trajectories):
policy_train — self-supervised training data for the resident-set policy:
structural features of branches in KVine's trajectory tree, labelled with
whether the branch was revisited within the next 5 steps.
benchmark_results.json — the full A / B / C benchmark output (cumulative
prefill FLOPs, TTFT, recomputed tokens… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/kvine-agent-trajectory-bench.crowd-anonymity-sets
Crowd — anonymity-set sizes for attribute combinations
Read this first
This dataset describes nobody. Every row is generated from a probability
model over attribute values built from published aggregate statistics. There
is no person in it, no record to link, and no index that could be searched for
an individual. The prose is synthetic.
The label is a head-count, not an identity. Each row's target is
log10(number of people in the reference population matching… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/crowd-anonymity-sets.assay-receipts
Assay Receipt Corpus
Signed internal receipts from an inference provider that is sometimes cheating, together
with the verdict an auditor reached on each one and the ground truth of which model actually
served the request.
Each row is a real receipt, not a summary statistic: it carries the prompt and output token
ids, the JL-projected sketch of the provider's hidden_states, the sign/rank invariants, and
an HMAC signature. With the gpt2 weights you can recompute the sketch… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/assay-receipts.molt-benchmark-results
Molt · elastic on-device inference measurements
Everything measured while building Molt, a
runtime that moves a running generation onto a smaller model between two
tokens, carrying the KV cache across, so an on-device LLM under memory pressure
is neither reclaimed by the OS nor restarted from the prompt.
Published so the claims can be checked rather than taken on trust. The figures in
the repo README and the results page are generated from these files; nothing is
transcribed by… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/molt-benchmark-results.clearance-bench
Clearance Benchmark
A fully synthetic enterprise corpus for measuring how well a retrieval system
contains information flow: heavily overlapping ACLs, sensitivity levels,
time-limited grants, a grant/revoke timeline, and evaluation queries.
Generated by clearance.synth with seed 7. No real
documents, no real access-control lists, and no real identities. Regenerating
with the same seed reproduces this dataset byte for byte.
Why this exists
Query-time ACL filtering… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/clearance-bench.sludge-ui-counterfactuals
Sludge counterfactual UI corpus
This model does not determine legality.
It reports provisions that may be implicated and the screen elements that are the
factual basis for looking at them. Whether a provision is actually engaged depends on facts
no UI tree contains — the purposes of processing, the legal basis relied on, the audience, the
rest of the journey, prior consent, sector rules — and is an assessment for a qualified human.
It has no feature that labels a… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/sludge-ui-counterfactuals.halfword-bench
Halfword benchmark: conversational text under access-method cost models
This dataset pairs public conversational sentences with timing cost models for AAC
access methods, so that a prediction system can be scored in seconds to utterance
rather than in keystrokes saved.
It contains no data from AAC users. It is public conversational text plus simulation.
Configurations
utterances (12565 rows) -- normalised sentences with history,
pseudo-speaker, source and that… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/halfword-bench.claimcheck-eval
ClaimCheck Evaluation Set
158 hand-checked cases for claim-level groundedness verification: given an
answer and the context that was supplied to the model, is each specific claim in
that answer supported, derived, approximate, unsupported or contradicted?
Bilingual (87 English / 71 Japanese). Every case was run through
ClaimCheck v1.1.0, and the
dataset records what the tool actually returned, not just what it should
have returned.
Cases
158
Languages
English 87 ·… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/claimcheck-eval.scribe-koyobun-usage
Scribe usage-judgment dataset
Span-level data for judging context-dependent kanji/kana usage in Japanese official writing.
Generated by scripts/build_dataset.py in the GitHub repo.
What the claim rests on
The center of this dataset is the hard split: occurrences of words that appear in both usages
(kana and kanji) across the corpus. Because uniform dictionary replacement collapses a word to a single
spelling, it is structurally forced to mislabel one side of this… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/scribe-koyobun-usage.litmus-kernels
Litmus Kernel Verification Corpus
Correct and deliberately-broken Triton kernels, each broken one shipped with
the input that exposes it.
The corpus exists to measure one thing: how much of what a fixed-shape
torch.rand() allclose test calls "correct" actually is. On this corpus the
answer is that 88% of the planted bugs pass that test.
Columns
column
meaning
name
kernel identifier
family
elementwise / reduction / softmax / layernorm / matmul /… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/litmus-kernels.amazon-laptop-product-catalogdendro-lowbackground
Dendro low-background corpus
5034 arXiv records annotated with archival evidence of when they existed, produced by
Dendro v0.1.0.
2500 of them (49.7%) are low-background: an independent registration record
places them before 2021-01-01, i.e. before large-scale text generation. The name is from
metallurgy — low-background steel is steel smelted before the 1945 atmospheric tests: not
special steel, just ordinary steel that happens to predate the contamination, and valuable
because… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/dendro-lowbackground.dissent-synthetic-clause-ambiguity
Dissent — synthetic financial clauses with seeded formalization ambiguity
Code · Interactive Space
Every clause in this dataset is synthetic. No real contract, client document or
third-party text was used, quoted or paraphrased. The language follows standard market
forms of drafting so that it is representative of the constructions that cause real
formalization disputes.
What this is for
Autoformalization research is crowded at the producing end and empty at the… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/dissent-synthetic-clause-ambiguity.amazon-laptop-reviews-enriched
Dataset Card for "amazon-laptop-reviews-enriched"
More Information needed
MidjourneySrefsFiltered subset of deepghs/midjourney_captioned_23m_full including only the prompts with the sref param passed in midjourney.
groundhog-world-tapes
Groundhog world tapes -- groundhog/research-publish-v1
Recorded exogenous randomness for a deliberately non-deterministic agent environment.
Each row is one intercepted interaction with the outside world: a clock reading, an
entropy draw, a UUID, an HTTP response, a file mtime, or an environment-variable read.
These tapes replay with no API key, no server and no network. That is checked before
publication: the mock world is shut down and every tape is replayed against a closed… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/groundhog-world-tapes.ESGdatasetsESGDatasetamazon-laptop-gpu-reviewscairn-departure-return
cairn-departure-return
Departure -> return trajectories with exact ground truth, for measuring whether a
video world model keeps a world when the camera looks away.
Each episode: a room of well-separated but deliberately confusable objects (two
near-duplicate colour pairs, repeated shapes), and a camera that observes a target
object, turns away for a measured number of frames, and comes back from a
different viewpoint -- so a model cannot pass by replaying its last frame.
A… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/cairn-departure-return.ingot-chrono
Ingot — Chrono
Natural-language time expressions to an iCalendar RRULE + ISO-8601 start + IANA timezone
exception rules, as strict JSON.
Every label in this dataset was constructed before its sentence existed. A schedule object is
generated from an integer seed, then rendered into prose. No model, judge or annotator ever
decided what the answer was, so the label cannot be wrong -- it is the input to the pipeline.
Splits
split
rows
verified
easy / medium /… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/ingot-chrono.
