datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
funes-handoff-recall-benchmark
handover-vs-recall
A long investigation bloats an agent session until each new turn costs more to carry the context than to
do the work. Switching to a fresh session avoids that — but the findings have to travel somehow, and the
ways of moving them differ in cost. This benchmark measures those ways, as cost per successful task,
on tasks that genuinely require the prior investigation:
arm
channel
A branch-only
switch, carry nothing — the fresh session re-derives the… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-handoff-recall-benchmark.VisualWebInstruct-Recall
Introduction
This is the dataset recalled from Google Search from the seed images.
Links
Github|
Paper|
Website
Citation
@article{visualwebinstruct,
title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search},
author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu},
journal={arXiv preprint arXiv:2503.10582},
year={2025}
}
VSI-SUPER-Recall
VSI-SUPER-Recall
Website | Paper | GitHub | Models
Authors: Shusheng Yang*, Jihan Yang*, Pinzhi Huang†, Ellis Brown†, et al.
VSI-SUPER-Recall is a benchmark for testing long-horizon spatial observation and recall in arbitrarily long videos. It evaluates whether models can remember and recall the order in which unusual objects appeared across extended video sequences.
Overview
VSI-SUPER-Recall challenges models to:
Track object appearances across long videos (10-240… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/VSI-SUPER-Recall.RecaLLM-data
RecaLLM Training and Evaluation Data
Training and evaluation datasets for RecaLLM. Contains GRPO reinforcement learning training data (20K examples) and evaluation data across 7 context lengths (4K-128K tokens).
Datasets generated using the code in recallm/datasets/ — see there for generation scripts and augmentation details.
Usage
from datasets import load_dataset
# Load training data for a specific dataset
ds = load_dataset("kswhitecross/RecaLLM-data", "hotpotqa"… See the full description on the dataset page: https://huggingface.co/datasets/kswhitecross/RecaLLM-data.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 5.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 1.3000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think.trace-cheating-recall-500
Trace Cheating Recall 500
This dataset contains 500 SWE-agent traces selected to evaluate whether an LLM judge
detects observable solution leakage. It is the public data source for the
trace-cheating-recall-500 Prime environment.
The examples were derived from
PrimeIntellect/int4-syn-gen-swe-glm53-bash-2026-09-02
at revision 0e7a9ecddce8de9ea8f8c369b2dd39411d6dee7a.
Composition
500 unique traces, all labeled CHEATING
250 internet-retrieval cases
250 Git-history… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/trace-cheating-recall-500.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 4.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think.recalldb-product-recalls-sample
RecallDB — U.S. Product Recall Database (Sample)
Full dataset: recalldb.dataengineered.io · $49 one-time snapshot → Buy on Stripe · the same sample on Kaggle
127,783 official recalls · 292,790 recalled products · CPSC · FDA · FSIS · NHTSA · USCG · 100% source-linked
RecallDB is a normalized, provenance-tracked dataset of official U.S. federal product recalls. It joins five official source families into one relational model: CPSC consumer products, NHTSA vehicles, FDA/openFDA… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/recalldb-product-recalls-sample.kids-product-recalls
KindlyKiddo Kids’ Product Recall Data
Machine-readable CPSC, NHTSA, and FDA recall records for products made for or primarily used by children. The dataset is normalized by KindlyKiddo and refreshed weekly from official U.S. federal sources.
Files
kids-product-recalls.csv — flat tabular dataset used by the Hugging Face viewer.
kids-product-recalls.json — records plus refresh time, coverage, source, license, and repository metadata.
data-dictionary.csv — field… See the full description on the dataset page: https://huggingface.co/datasets/kindlykiddo/kids-product-recalls.Instruction_recall_dataset
CanaryBench-PII
Frequency-aware canary injection benchmark for auditing memorization
in finetuned language models, built on the AI4Privacy PII reconstruction
task.
Dataset Description
This dataset is part of CanaryBench, a benchmark for evaluating
memorization in finetuned language models across repetition tiers
and privacy regimes.
Frequency tiers: 1×, 10×, 50×
PII types: EMAIL, PHONE
Member canaries: 770
Reference canaries: 1000
Tasks: PII detection, secret… See the full description on the dataset page: https://huggingface.co/datasets/anony-mouse123/Instruction_recall_dataset.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 2.2000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think.vett-cpsc-recalls
Vett CPSC Product Recall Corpus
Normalized U.S. Consumer Product Safety Commission (CPSC) product recall records,
refreshed periodically from CPSC's live recall feed.
Source: U.S. Consumer Product Safety Commission (cpsc.gov). As a work of the U.S.
federal government, the underlying data is in the public domain under 17 U.S.C. Section 105,
not subject to copyright. This normalization/compilation is provided by
Vett (Wimberly Solutions LLC).
Fields: recall_id, source… See the full description on the dataset page: https://huggingface.co/datasets/Wim-Sol/vett-cpsc-recalls.nhtsa-vehicle-recalls
NHTSA Vehicle Recalls 1966–March 2026
One-time dated snapshot.
Recall campaign and affected-product records with defect and remedy text.
Verified coverage: Recall dates: 1966-01-19 through 2026-03-26.
Records: 176,073 product rows; 29,865 distinct campaign IDs.
This repository contains a 1,000-row public sample, not the full package. A sample does not establish complete historical coverage.
Limitations
Multiple product/model rows can belong to the same campaign.… See the full description on the dataset page: https://huggingface.co/datasets/claritystorm/nhtsa-vehicle-recalls.funes-recall-session-pi-traces
dacorvo/funes-recall-session-pi-traces
pi coding-agent session traces produced by
agentcap runs. Each run
contributes one folder under data/<run_id>/; inside, one file per
session in pi's native export format.
The on-the-wire HTTP captures for these same runs live in
dacorvo/funes-recall-session-captures.
Both belong to the
funes-recall-session Collection
— join on run_id to align captures with traces.
recall_test_datasetbiobert-ner-fda-recalls-dataset
Dataset Card for FDA CDRH Device Recalls NER Dataset
This is a FDA Medical Device Recalls Dataset Created for Medical Device Named Entity Recognition (NER)
Dataset Details
Dataset Description
This dataset was created for the purpose of performing NER tasks.
It utilizes the OpenFDA Device Recalls dataset, which has been processed and annotated for performing NER.
The Device Recalls dataset has been further processed to extract the recall action element, which… See the full description on the dataset page: https://huggingface.co/datasets/mfarrington/biobert-ner-fda-recalls-dataset.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.05.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.01.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.1.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-10m-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-10m-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 10M notes + note-conditioned trajectory mixture, and no KL regularization.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-10m-think.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-30m-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-30m-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 30M notes + note-conditioned trajectory mixture, and no KL regularization.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-30m-think.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and no KL regularization.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think.recalls-by-barcode
ProductGuru — Consumer Product Recalls by Barcode (EAN/GTIN)
36,654 recall notices across 32,668 distinct retail barcodes, compiled from
11 government registers. One row per (barcode, authority, notice).
100% of rows link to the issuing authority's own notice.
Why this exists
Most official recall registers do not publish a barcode. Across the registers we mirror, only
29.3% of consumer-product recalls carry one — France's RappelConso publishes a barcode on
100% of… See the full description on the dataset page: https://huggingface.co/datasets/PROGU2026/recalls-by-barcode.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-1m-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-1m-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 1M notes + note-conditioned trajectory mixture, and no KL regularization.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-1m-think.ppb-kenya-recalls-dataset
PPB Kenya Recalls
Structured data on medicine and medical device recalls and rapid alerts issued by the Pharmacy and Poisons Board (PPB) of Kenya, covering 2016 to 2025.
Load the dataset
from datasets import load_dataset
# Recall notices (default)
recalls = load_dataset("afiadata/ppb-kenya-recalls", "recalls")
# Rapid alerts
alerts = load_dataset("afiadata/ppb-kenya-recalls", "rapid_alerts")
Or with pandas directly:
import pandas as pd
recalls =… See the full description on the dataset page: https://huggingface.co/datasets/AfiadataKe/ppb-kenya-recalls-dataset.medmcqa-bonvoyage-39995_recall_0.5-qwen3-4b-medical-qa-htrecall-sessions
author-samzong_project-recall_time-2026-09-12
Local AI coding sessions from the Recall project, exported and redacted with Recall and published by samzong.
Selection
Window: 2026-09-12T00:00:00+00:00 to 2026-09-13T00:00:00+00:00 on session.started_at
Sessions: 1
Sources: all
Thread roles: all
Files
author-samzong_project-recall_time-2026-09-12.recall.jsonl — one JSON object per session, Recall export schema version 7
manifest.json — selection… See the full description on the dataset page: https://huggingface.co/datasets/samzong/recall-sessions.bergson-recall-9000recall-rewrite-oasst1
Recall Rewrite OASST1: knowledge-aligned SFT data
Data release for the paper "Stick to What You Know: A Study of Knowledge-Aligned Supervised Fine-Tuning"
(Becker, Kemmler, Thulke, Schäfer, Dugast, Ney; accepted at EMNLP 2026, Main Conference).
Knowledge-aligned SFT constrains supervised fine-tuning targets to what the base model already knows.
Recall Rewrite implements this without external evidence: every gold response of the SFT set is
decomposed into atomic claims, each… See the full description on the dataset page: https://huggingface.co/datasets/apptek-com/recall-rewrite-oasst1.
