datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-runtime-recovery-bench
Agent Runtime Recovery Benchmark
This fully public dataset combines the current observation-restricted Agent
Runtime Recovery Benchmark with causally qualified native runtime cases for real
coding agents.
Contents
Source
Cases
Description
benchmark
769
Observation-restricted runtime recovery cases
real_agent_native
125
90 OpenHands and 35 mini-SWE-agent native runtime cases
Total
894
One unified public dataset
All rows are stored in one all… See the full description on the dataset page: https://huggingface.co/datasets/zitong1/agent-runtime-recovery-bench.NovGauge
NovGauge
NovGauge evaluates paper similarity along three dimensions: task, problem,
and method. This dataset contains final benchmark labels and bibliographic
metadata for pairwise classification and multi-paper grouping.
Dataset repository: ZitaGo/NovGauge.
Release contents
File
Records
Contents
data/positives.json
463
Paper pairs similar in at least one annotated dimension
data/negatives.json
156
Paper pairs dissimilar in the specified annotated… See the full description on the dataset page: https://huggingface.co/datasets/ZitaGo/NovGauge.
