datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
funes-handoff-recall-benchmark
handover-vs-recall
A long investigation bloats an agent session until each new turn costs more to carry the context than to
do the work. Switching to a fresh session avoids that — but the findings have to travel somehow, and the
ways of moving them differ in cost. This benchmark measures those ways, as cost per successful task,
on tasks that genuinely require the prior investigation:
arm
channel
A branch-only
switch, carry nothing — the fresh session re-derives the… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-handoff-recall-benchmark.LongVideo-Reason-4k-Video-Crop-Handoff-20260911
LongVideo-Reason 4k · Video Crop 合成移交包
公开仓库,文件访问需要人工审批。 只有仓库根目录出现 READY.json 且 complete=true 时,才表示所有 QA、视频、pipeline 和校验信息已齐备;此前为准备/上传阶段。
本包用于将原视频和原始 QA 重新合成为视频工具轨迹。它不是已经审核通过的 SFT 数据,也不把原论文 reasoning 当作工具轨迹监督。
内容
文件
用途
data/qa.jsonl
4,000 条原始 LongVideo-Reason train QA、原选项、原答案和来源
videos/*.mp4
配套原视频;与 QA 的 video_path 对应
data/video_manifest.jsonl
每个视频的 SHA-256、CRC、ffprobe 时长、尺寸和镜像来源
data/selection_report.json
最终数量、时长分布、去重和筛选范围… See the full description on the dataset page: https://huggingface.co/datasets/b1intern/LongVideo-Reason-4k-Video-Crop-Handoff-20260911.adaption-piguard-translation-handoff
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-piguard_translation_handoff
This dataset contains 100 English prompts curated for translation guardrail research, comprising a balanced mix of 50 benign instructions and 50 prompt injection attempts. The samples include diverse content such as jailbreak personas, requests for harmful actions like spyware installation, and complex instruction overrides. Each entry is… See the full description on the dataset page: https://huggingface.co/datasets/melanieyes/adaption-piguard-translation-handoff.AGENT-HANDOFF-REPRO-01arena-pokerkit-hands
Arena PokerKit Hands (S8 archive — offline practice data)
A real-hand dataset of 6-max No-Limit Texas Hold'em poker played by 19 AI agents
on the dev.fun Arena Beta — frontier LLMs, OSS solver
bots, and equity heuristics — captured during the S8 benchmark run
(May 5–6, 2026).
What this is and is not. This is a derived, settled-hand archive in a
normalized snake_case schema. It is not a mirror of the live
/texas/benchmark/status.table (or /texas/pending-actions[].) shape. Use
it for… See the full description on the dataset page: https://huggingface.co/datasets/dannyobito/arena-pokerkit-hands.paper-evo-hand
