CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 5.0000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think.tabulartext-generation1K<n<10K0 likes291 downloads3d agoHugging Face02violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 1.3000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think.tabulartext-generation1K<n<10K0 likes274 downloads3d agoHugging Face03violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 4.0000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think.tabulartext-generation1K<n<10K0 likes272 downloads3d agoHugging Face04anony-mouse123 /Instruction_recall_dataset CanaryBench-PII Frequency-aware canary injection benchmark for auditing memorization in finetuned language models, built on the AI4Privacy PII reconstruction task. Dataset Description This dataset is part of CanaryBench, a benchmark for evaluating memorization in finetuned language models across repetition tiers and privacy regimes. Frequency tiers: 1×, 10×, 50× PII types: EMAIL, PHONE Member canaries: 770 Reference canaries: 1000 Tasks: PII detection, secret… See the full description on the dataset page: https://huggingface.co/datasets/anony-mouse123/Instruction_recall_dataset.texttext-generation10K<n<100K0 likes107 downloads2mo agoHugging Face05violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 2.2000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think.tabulartext-generation1K<n<10K0 likes97 downloads3d agoHugging Face06violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.05. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think.tabulartext-generation1K<n<10K0 likes75 downloads1d agoHugging Face07violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.01. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think.tabulartext-generation1K<n<10K0 likes73 downloads1d agoHugging Face08violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.1. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think.tabulartext-generation1K<n<10K0 likes73 downloads1d agoHugging Face09violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-10m-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-10m-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 10M notes + note-conditioned trajectory mixture, and no KL regularization. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-10m-think.tabulartext-generation1K<n<10K0 likes71 downloads1d agoHugging Face10violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-30m-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-30m-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 30M notes + note-conditioned trajectory mixture, and no KL regularization. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-30m-think.tabulartext-generation1K<n<10K0 likes71 downloads1d agoHugging Face11violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and no KL regularization. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think.tabulartext-generation1K<n<10K0 likes70 downloads1d agoHugging Face12violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-1m-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-1m-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 1M notes + note-conditioned trajectory mixture, and no KL regularization. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-1m-think.tabulartext-generation1K<n<10K0 likes65 downloads1d agoHugging Face13samzong /recall-sessions author-samzong_project-recall_time-2026-09-12 Local AI coding sessions from the Recall project, exported and redacted with Recall and published by samzong. Selection Window: 2026-09-12T00:00:00+00:00 to 2026-09-13T00:00:00+00:00 on session.started_at Sessions: 1 Sources: all Thread roles: all Files author-samzong_project-recall_time-2026-09-12.recall.jsonl — one JSON object per session, Recall export schema version 7 manifest.json — selection… See the full description on the dataset page: https://huggingface.co/datasets/samzong/recall-sessions.texttext-generationn<1K0 likes52 downloads6d agoHugging Face14apptek-com /recall-rewrite-oasst1 Recall Rewrite OASST1: knowledge-aligned SFT data Data release for the paper "Stick to What You Know: A Study of Knowledge-Aligned Supervised Fine-Tuning" (Becker, Kemmler, Thulke, Schäfer, Dugast, Ney; accepted at EMNLP 2026, Main Conference). Knowledge-aligned SFT constrains supervised fine-tuning targets to what the base model already knows. Recall Rewrite implements this without external evidence: every gold response of the SFT set is decomposed into atomic claims, each… See the full description on the dataset page: https://huggingface.co/datasets/apptek-com/recall-rewrite-oasst1.tabulartext-generation10K<n<100K0 likes46 downloads25d agoHugging Face15cds-jb /qwen3-8b-codi-multihop-recall-data CODI training data — multi-hop recall & pointer-chase (single-token-node reasoning) The training data + generators + load-bearing eval code for two Qwen3-8B CODI latent-reasoning organisms: cds-jb/qwen3-8b-codi-multihop-recall and cds-jb/qwen3-8b-codi-pointer-chase. Both tasks are single-token-node serial-reasoning problems: every intermediate and the final answer is a single token (in both the Qwen3 and Gemma3 tokenizers), so each CODI latent can in principle be read with a… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/qwen3-8b-codi-multihop-recall-data.text-generation0 likes39 downloads3mo agoHugging Face16Ghostgim /cybersec-fact-recall Cybersec Fact-Recall Benchmark (GhostLM v2) Free-form short-answer benchmark for small cybersecurity language models. Built and used by the GhostLM project as the truth metric for the ghost-base v1.0 acceptance gate. Why this exists Multiple-choice cybersec benchmarks like CTIBench and SecQA reward register matching (the model picks the option that "looks like" a security answer) as much as actual factual recall. A small from- scratch model can hit 28-30% on those without… See the full description on the dataset page: https://huggingface.co/datasets/Ghostgim/cybersec-fact-recall.texttext-generationn<1K0 likes29 downloads5mo agoHugging Face17ping98k /squad-qwq-recall-1k squad-qwq-recall-1k This dataset is planned to be used as SFT to create the recall-writer model in flow step 1 Recall Writer Flow Purpose: To synthetically augment data for training a smaller model to effectively recall new knowledge from documents and apply it in the thinking process. Step 1: Distill recall traces from the reasoning model Ask the reasoning model to recall memories related to the question. The expectation is that the reasoning model, trained… See the full description on the dataset page: https://huggingface.co/datasets/ping98k/squad-qwq-recall-1k.texttext-generation1K<n<10K1 likes17 downloads1y agoHugging Face18Mandotosh /risk-routed-kv-exact-recall-benchmark Risk-Routed KV Exact-Recall Benchmark This dataset contains controlled synthetic exact-recall examples used to evaluate risk-routed heterogeneous KV memory policies for long-context Transformer inference. The benchmark is designed for testing whether a model can retrieve exact strings from long contexts under different KV-cache policies: Full KV Uniform low-bit Quantized KV Risk-routed heterogeneous KV, where exact-critical spans stay in Full KV and background context is… See the full description on the dataset page: https://huggingface.co/datasets/Mandotosh/risk-routed-kv-exact-recall-benchmark.texttext-generationn<1K1 likes16 downloads2mo agoHugging Face19schoggie /java-agentic-recall-en Java Agentic + Recall (English) Synthetic training data for fine-tuning a Java-specialist agentic coding model with explicit long-context recall capability. Companion dataset to a Qwen3.6-35B-A3B QLoRA SFT pilot. Composition Split Source Rows train DeepSeek V4 Pro (synthetic agentic Java traces) 3873 train Synthetic positional recall — short context (~26K tok) 120 train Synthetic positional recall — long context (50K-180K tok) 46 train total 4039 eval… See the full description on the dataset page: https://huggingface.co/datasets/schoggie/java-agentic-recall-en.text-generation1K<n<10K0 likes15 downloads4mo agoHugging Face20violetxi /wmrl-v4-recall-trajectories Recall trajectories (WM-RL v4) 19,481 rewritten agent trajectories over a synthetic law-firm document-management system. Each row is one recall trajectory: an original agent rollout in which every tool call and every tool observation is byte-identical to the original run, and only the model's thinking traces were regenerated. The regenerated narration is therefore post-hoc recall of a record the model can no longer see — which is exactly the signal these were built to train and… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/wmrl-v4-recall-trajectories.tabulartext-generation10K<n<100K0 likes12 downloads13d agoHugging Face21endsky /sera-4.5-django-t2-recall05-toolcallsgated SERA-4.5A Django T2 (Recall=0.5) Toolcalls This dataset contains normalized multi-turn tool-calling trajectories derived from: Source dataset: allenai/Sera-4.5A-Django-T2 Filter: line_level_recall == 0.5 Splits train.jsonl: 6200 records val.jsonl: 331 records Format Each line is a JSON object with: id: trajectory id messages: normalized chat/tool-call messages metadata: includes instance_id, func_name, func_path, line_level_recall Processing… See the full description on the dataset page: https://huggingface.co/datasets/endsky/sera-4.5-django-t2-recall05-toolcalls.texttext-generation1K<n<10K0 likes3 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.