Jack-Jieke-Wu/Paper-Reviewing-Exam-Trails
Paper-Reviewing-Exam-Trails Review-run trail archive for the Paper-Reviewing-Exam benchmark. Each trail captures one agent review run: the submitted review.md and review.json, the agent's brain/, manifests, logs, and verifier output, so a human expert can assess the review against the exact materials it was written from. Trails contain review content and are public. Review text can be read from this dataset; uploading a trail publishes it. Confirm you accept that exposure before… See the full description on the dataset page: https://huggingface.co/datasets/Jack-Jieke-Wu/Paper-Reviewing-Exam-Trails.
0133
