datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anamnesis-bench
AnamnesisBench
AnamnesisBench is an evaluation benchmark for numerical reliability in LLM research agents.
It focuses on a practical failure mode: an agent writes or accepts a financial research artifact that
looks plausible, but contains a wrong, unsupported, or misattributed number.
The benchmark is not intended as training data. It is a set of test cases, source packets, expected
truth values, and deterministic scoring scripts. You run your own model or verifier, then score… See the full description on the dataset page: https://huggingface.co/datasets/pppop7/anamnesis-bench.video_agent_trace
Friday prompt delivery — 30K text-only corpus
本目录是可直接消费的 prompt 交付包;所有内容均来自已完成的 Friday GPT 文本生产,不包含视频,也不表示 H3/视频评测通过。
文件
文件
内容
top_historical_test_prompts_4of5.jsonl
第一轮多轮历史实验中全局最高接受率的两条 test prompt(各 4/5)。
family_champion_test_prompts.jsonl
四个家族各自采用的 champion test prompt;包含 4/5、4/5、3/5、3/5 的历史结果与选择理由。
base_prompts_1200.jsonl
1,200 条干净 base prompt,每条有稳定 base_id。
injection_prompts_30000.jsonl
30,000 条最终 injection prompt;prompt 是可直接读取的 prompt… See the full description on the dataset page: https://huggingface.co/datasets/ppppppz/video_agent_trace.
