CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nagarhimanshu37 /brain-memory 🧠 NIFTY AI Agent: Memory OS Cloud Snapshot Cloud backup repository for the NIFTY 50 Autonomous AI Agent Memory OS. • Repository: nagarhimanshu37/brain-memory• Total Stored Records: 235• Last Synchronized: 2026-09-25 17:46:58 UTC 📊 Partition Statistics Partition Records Description conversation_memory 83 Multi-turn trader dialogues & intent logs episodic_memory 50 Trading day episodes (facts vs interpretations) experience_memory 50 Crystallized… See the full description on the dataset page: https://huggingface.co/datasets/nagarhimanshu37/brain-memory.texttext-generationn<1K0 likes181 downloads20h agoHugging Face02NagaYu /isotope-bench Isotope Bench An indirect-prompt-injection benchmark for tool-calling agents, plus the complete audit trail of one recorded run: 438 influence certificates, one for every action an agent attempted across five defence conditions. Built for Isotope, which tracks untrusted influence inside the forward pass. The corpus is independent of that method and usable with any defence. 💻 Code: https://github.com/NagaYu/isotope 🤗 Demo: https://huggingface.co/spaces/NagaYu/isotope 🤗… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/isotope-bench.tabulartext-generationn<1K1 likes108 downloads20d agoHugging Face03NagaYu /mondegreen-asr-errors Mondegreen ASR error pairs (ASR hypothesis, gold text) pairs for Japanese ASR post-correction. This build is simulated -- errors come from a phonetic corruption model, not from a real ASR system. It exists so the whole pipeline (gate training, benchmarks, figures, CI) is reproducible without a GPU. Treat every number derived from it as a stated assumption, not a measurement. How it was made synthetic text -> phonetic corruption model (mondegreen.simulate) ->… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/mondegreen-asr-errors.automatic-speech-recognition1K<n<10K0 likes72 downloads1mo agoHugging Face04NagaYu /signpost-label-quality Signpost label quality Labels from accessibility trees, each one classed as good or as one of seven ways a label can fail to mean anything. It is built for the question that is left over after axe-core and Xcode's Accessibility Inspector have both passed: there is a name, but does the name identify this control? Repository: NagaYu/signpost-label-quality Code, evaluation, and the builder for this dataset https://github.com/NagaYu/signpost Model trained on it… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/signpost-label-quality.texttext-classification1K<n<10K0 likes54 downloads13d agoHugging Face05NagaYu /halfword-bench Halfword benchmark: conversational text under access-method cost models This dataset pairs public conversational sentences with timing cost models for AAC access methods, so that a prediction system can be scored in seconds to utterance rather than in keystrokes saved. It contains no data from AAC users. It is public conversational text plus simulation. Configurations utterances (12565 rows) -- normalised sentences with history, pseudo-speaker, source and that… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/halfword-bench.tabulartext-generation10K<n<100K0 likes47 downloads24d agoHugging Face06agnivamaiti /naganlp-conversational-corpus NagaNLP Conversational Corpus Dataset Summary This is the official conversational dataset for the NagaNLP project. It contains 10,021 instruction-following pairs (User / Assistant) in Nagamese (Naga Pidgin), intended for instruction-tuning or evaluating conversational models in a low-resource creole language. Supported Tasks Text Generation / Instruction Following: given a user prompt in Nagamese, generate an appropriate assistant response in… See the full description on the dataset page: https://huggingface.co/datasets/agnivamaiti/naganlp-conversational-corpus.text-generation10K<n<100K1 likes32 downloads3mo agoHugging Face07NagaYu /dendro-lowbackground Dendro low-background corpus 5034 arXiv records annotated with archival evidence of when they existed, produced by Dendro v0.1.0. 2500 of them (49.7%) are low-background: an independent registration record places them before 2021-01-01, i.e. before large-scale text generation. The name is from metallurgy — low-background steel is steel smelted before the 1945 atmospheric tests: not special steel, just ordinary steel that happens to predate the contamination, and valuable because… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/dendro-lowbackground.tabulartext-classification1K<n<10K0 likes24 downloads2mo agoHugging Face08NagaYu /ingot-chrono Ingot — Chrono Natural-language time expressions to an iCalendar RRULE + ISO-8601 start + IANA timezone exception rules, as strict JSON. Every label in this dataset was constructed before its sentence existed. A schedule object is generated from an integer seed, then rendered into prose. No model, judge or annotator ever decided what the answer was, so the label cannot be wrong -- it is the input to the pipeline. Splits split rows verified easy / medium /… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/ingot-chrono.tabulartext-generation10K<n<100K0 likes12 downloads1mo agoHugging Face09NagaYu /parity-fertility-atlas Parity fertility atlas How many tokens each tokenizer charges for the same meaning, measured on a parallel corpus (opus100). column meaning tokenizer_id the tokenizer measured lang ISO code tokens_per_char tokens per NFC character, excluding whitespace tokens_per_word tokens per whitespace word; null for scripts without word spaces parity_ratio tokens(target) / tokens(aligned English) — the headline parity_ratio_median median of the per-sentence ratios… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/parity-fertility-atlas.tabulartext-generationn<1K0 likes11 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.