aac
Datasets
All datasets matching “aac”aacr-bench-harbor
AACR-Bench Harbor
AACR-Bench Harbor packages AACR-Bench as independent Harbor tasks for code-review agents. Each task checks out one pull request at its pinned head commit, keeps the base as aacr-base, and grades a structured findings file with AACR-Bench's path, side, line, and semantic matching stages.
This repository is not a fork of AACR-Bench or Harbor.
Install
Requirements:
Python 3.12 or newer
uv
Git
Harbor 0.21.0
Docker for local task runs, or Harbor's… See the full description on the dataset page: https://huggingface.co/datasets/osolmaz/aacr-bench-harbor.aacr-bench
Dataset for Running AACR-Bench
English | 简体中文
This is a test set designed for automated code review reflection models, primarily aiming to evaluate the extent to which a model can intercept low-quality review comments. The dataset contains 2,145 code review comments, consisting of 1,505 expert-verified correct comments and 640 incorrect comments.
This data is part of the AACR-Bench project and is provided by the Alibaba Aone team.
Data Sample
Each sample in the… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Aone/aacr-bench.aac_redditThis dataset contains sentences from Reddit.
Each sentence is scored according to how similar it was to a spoken (dialogue_prob) or written (forum_prob) communication.
For details of how we scored the sentences, see our EMNLP 2025 paper.
However, we did not make use of this dataset in the results reported in this paper.
Our early experiments showed no gains versus using the much smaller C4 and Subtitle training sets.
aac_c4_deberta_classifiedThis dataset contains sentences from the Colossal Clean Crawled Corpus corpus.
Each sentence is scored according to how similar it was to a spoken (dialogue_prob) or written (forum_prob) communication.
See our EMNLP 2025 paper for details.
aac_c4_deberta_classified_0.90This dataset contains sentences from the Colossal Clean Crawled Corpus corpus.
This is a subset of the dataset figmtu/aac_c4_deberta_classified.
It contains only the sentences that had a dialogue or forum probability of 0.90 or greater.
See our EMNLP 2025 paper for details.
Ludus
🧠 Ludus (Loose Goosey Release)
This is the Ludus Archive. It is not tidy. It is not structured. But it is alive.
Inside, you’ll find dozens of .docx files uploaded directly—notes, scrolls, recursive fragments, ruptures, and ritual games. This version does not yet follow a clean split-manifest, but it contains everything that matters.
🌀 What’s Inside
Philosophical games of contradiction
Recursions between human and synthetic minds
Symbolic and ethical tests
The… See the full description on the dataset page: https://huggingface.co/datasets/AAC3322/Ludus.
