datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
brain-memory
🧠 NIFTY AI Agent: Memory OS Cloud Snapshot
Cloud backup repository for the NIFTY 50 Autonomous AI Agent Memory OS.
• Repository: nagarhimanshu37/brain-memory• Total Stored Records: 235• Last Synchronized: 2026-09-25 17:46:58 UTC
📊 Partition Statistics
Partition
Records
Description
conversation_memory
83
Multi-turn trader dialogues & intent logs
episodic_memory
50
Trading day episodes (facts vs interpretations)
experience_memory
50
Crystallized… See the full description on the dataset page: https://huggingface.co/datasets/nagarhimanshu37/brain-memory.isotope-bench
Isotope Bench
An indirect-prompt-injection benchmark for tool-calling agents, plus the
complete audit trail of one recorded run: 438 influence certificates, one for
every action an agent attempted across five defence conditions.
Built for Isotope, which tracks
untrusted influence inside the forward pass. The corpus is independent of that
method and usable with any defence.
💻 Code: https://github.com/NagaYu/isotope
🤗 Demo: https://huggingface.co/spaces/NagaYu/isotope
🤗… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/isotope-bench.signpost-label-quality
Signpost label quality
Labels from accessibility trees, each one classed as good or as one of seven
ways a label can fail to mean anything. It is built for the question that is
left over after axe-core and Xcode's Accessibility Inspector have both
passed: there is a name, but does the name identify this control?
Repository: NagaYu/signpost-label-quality
Code, evaluation, and the builder for this dataset
https://github.com/NagaYu/signpost
Model trained on it… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/signpost-label-quality.halfword-bench
Halfword benchmark: conversational text under access-method cost models
This dataset pairs public conversational sentences with timing cost models for AAC
access methods, so that a prediction system can be scored in seconds to utterance
rather than in keystrokes saved.
It contains no data from AAC users. It is public conversational text plus simulation.
Configurations
utterances (12565 rows) -- normalised sentences with history,
pseudo-speaker, source and that… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/halfword-bench.dendro-lowbackground
Dendro low-background corpus
5034 arXiv records annotated with archival evidence of when they existed, produced by
Dendro v0.1.0.
2500 of them (49.7%) are low-background: an independent registration record
places them before 2021-01-01, i.e. before large-scale text generation. The name is from
metallurgy — low-background steel is steel smelted before the 1945 atmospheric tests: not
special steel, just ordinary steel that happens to predate the contamination, and valuable
because… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/dendro-lowbackground.ingot-chrono
Ingot — Chrono
Natural-language time expressions to an iCalendar RRULE + ISO-8601 start + IANA timezone
exception rules, as strict JSON.
Every label in this dataset was constructed before its sentence existed. A schedule object is
generated from an integer seed, then rendered into prose. No model, judge or annotator ever
decided what the answer was, so the label cannot be wrong -- it is the input to the pipeline.
Splits
split
rows
verified
easy / medium /… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/ingot-chrono.parity-fertility-atlas
Parity fertility atlas
How many tokens each tokenizer charges for the same meaning, measured on a
parallel corpus (opus100).
column
meaning
tokenizer_id
the tokenizer measured
lang
ISO code
tokens_per_char
tokens per NFC character, excluding whitespace
tokens_per_word
tokens per whitespace word; null for scripts without word spaces
parity_ratio
tokens(target) / tokens(aligned English) — the headline
parity_ratio_median
median of the per-sentence ratios… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/parity-fertility-atlas.
