recall
Datasets
All datasets matching “recall”funes-handoff-recall-benchmark
handover-vs-recall
A long investigation bloats an agent session until each new turn costs more to carry the context than to
do the work. Switching to a fresh session avoids that — but the findings have to travel somehow, and the
ways of moving them differ in cost. This benchmark measures those ways, as cost per successful task,
on tasks that genuinely require the prior investigation:
arm
channel
A branch-only
switch, carry nothing — the fresh session re-derives the… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-handoff-recall-benchmark.VisualWebInstruct-Recall
Introduction
This is the dataset recalled from Google Search from the seed images.
Links
Github|
Paper|
Website
Citation
@article{visualwebinstruct,
title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search},
author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu},
journal={arXiv preprint arXiv:2503.10582},
year={2025}
}
VSI-SUPER-Recall
VSI-SUPER-Recall
Website | Paper | GitHub | Models
Authors: Shusheng Yang*, Jihan Yang*, Pinzhi Huang†, Ellis Brown†, et al.
VSI-SUPER-Recall is a benchmark for testing long-horizon spatial observation and recall in arbitrarily long videos. It evaluates whether models can remember and recall the order in which unusual objects appeared across extended video sequences.
Overview
VSI-SUPER-Recall challenges models to:
Track object appearances across long videos (10-240… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/VSI-SUPER-Recall.Redmond-Sentence-Recall
Dataset Summary
The Redmond Sentence Recall (RSR) measures a child’s ability to repeat sentences that contain regular past tense forms and past participle forms (e.g., “He kicked” vs. “He was kicked”). This task helps identify language impairments, with each child repeating 16 sentences heard through headphones. The dataset includes anonymized audio recordings of these repetitions.
What makes the RSR dataset uniquely valuable is its focus on sentence recall using both regular past… See the full description on the dataset page: https://huggingface.co/datasets/ai4exceptionaled/Redmond-Sentence-Recall.RecaLLM-data
RecaLLM Training and Evaluation Data
Training and evaluation datasets for RecaLLM. Contains GRPO reinforcement learning training data (20K examples) and evaluation data across 7 context lengths (4K-128K tokens).
Datasets generated using the code in recallm/datasets/ — see there for generation scripts and augmentation details.
Usage
from datasets import load_dataset
# Load training data for a specific dataset
ds = load_dataset("kswhitecross/RecaLLM-data", "hotpotqa"… See the full description on the dataset page: https://huggingface.co/datasets/kswhitecross/RecaLLM-data.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 5.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think.
