datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pyaptamer-AptaComapt-eval
🚨 APT-Eval Dataset 🚨
Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing
📝 Paper, 🖥️ Github, 🎥 Recording
This repository contains the official dataset of the ACL 2025 paper 'Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing'
APT-Eval is the first and largest dataset to evaluate the AI-text detectors behavior for AI-polished texts.
It contains almost 15K text samples, polished by 5 different LLMs, for 6 different domains, with 2 major… See the full description on the dataset page: https://huggingface.co/datasets/smksaha/apt-eval.AptaBench_dataset
AptaBench
AptaBench is a benchmark for aptamer–small-molecule interaction prediction. It contains curated DNA/RNA aptamer–ligand pairs with standardized sequences, canonical SMILES, experimentally grounded active/inactive labels, quantitative affinity values where available, and fixed leakage-aware evaluation splits.
This repository is provided for anonymous peer review. Author identities, affiliations, acknowledgements, citation information, and non-anonymous project… See the full description on the dataset page: https://huggingface.co/datasets/aptabench-anonymous/AptaBench_dataset.AptaBench_dataset
AptaBench
AptaBench is a benchmark for aptamer–small-molecule interaction prediction. It contains curated DNA/RNA aptamer–ligand pairs with standardized sequences, canonical SMILES, experimentally grounded active/inactive labels, quantitative affinity values where available, and fixed leakage-aware evaluation splits.
This repository is provided for anonymous peer review. Author identities, affiliations, acknowledgements, citation information, and non-anonymous project… See the full description on the dataset page: https://huggingface.co/datasets/swampfireee/AptaBench_dataset.apthttps://github.com/Advancing-Machine-Human-Reasoning-Lab/apt
@inproceedings{nighojkar-licato-2021-improving,
title = "Improving Paraphrase Detection with the Adversarial Paraphrasing Task",
author = "Nighojkar, Animesh and
Licato, John",
booktitle = "Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)",
month = aug,
year = "2021"… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/apt.APTO-HarmBench-JA
APTO-HarmBench-JA
APTO-HarmBench-JA is a Japanese translated and annotated version of the HarmBench dataset for AI safety evaluation research.
HarmBench is a standardized evaluation framework for automated red teaming and robust refusal. It is designed to evaluate whether language models can appropriately refuse harmful requests across a wide range of risk categories.
This dataset includes:
Japanese translations of HarmBench prompts
Japanese refusal responses
Refusal reasoning… See the full description on the dataset page: https://huggingface.co/datasets/APTO-001/APTO-HarmBench-JA.AptaComColumns:
Aptamer Name;
Aptamer Sequence;
Reference (DOI/PUBMED ID);
Target Name;
Target Sequence;
External id (PDB/ATCC/PUBCHEM);
Original database;
APTO-XSTest-JA
APTO-XSTest-JA
APTO-XSTest-JA is a Japanese translated and annotated version of the XSTest dataset for AI safety evaluation research.
XSTest is a test suite designed to identify exaggerated safety behaviours in large language models, including cases where models refuse clearly safe prompts because they contain sensitive wording or resemble unsafe requests.
This dataset includes:
Japanese translations of XSTest prompts
Japanese refusal / non-refusal reference responses
Refusal… See the full description on the dataset page: https://huggingface.co/datasets/APTO-001/APTO-XSTest-JA.APTO-SorryBench-JA
APTO-SorryBench-JA
APTO-SorryBench-JA is a Japanese translated and annotated version of the original Sorry-Bench dataset for AI safety evaluation research.
Sorry-Bench is a benchmark designed to evaluate whether Large Language Models (LLMs) appropriately refuse harmful requests while still providing helpful responses to safe requests.
The original English prompts are preserved alongside the Japanese translations to improve traceability and facilitate comparison with the original… See the full description on the dataset page: https://huggingface.co/datasets/APTO-001/APTO-SorryBench-JA.tokenized_ds_stats_apt4
