datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bybit-linear-perps-aptusdtapt-eval
🚨 APT-Eval Dataset 🚨
Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing
📝 Paper, 🖥️ Github, 🎥 Recording
This repository contains the official dataset of the ACL 2025 paper 'Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing'
APT-Eval is the first and largest dataset to evaluate the AI-text detectors behavior for AI-polished texts.
It contains almost 15K text samples, polished by 5 different LLMs, for 6 different domains, with 2 major… See the full description on the dataset page: https://huggingface.co/datasets/smksaha/apt-eval.AptaBench_dataset
AptaBench
AptaBench is a benchmark for aptamer–small-molecule interaction prediction. It contains curated DNA/RNA aptamer–ligand pairs with standardized sequences, canonical SMILES, experimentally grounded active/inactive labels, quantitative affinity values where available, and fixed leakage-aware evaluation splits.
This repository is provided for anonymous peer review. Author identities, affiliations, acknowledgements, citation information, and non-anonymous project… See the full description on the dataset page: https://huggingface.co/datasets/aptabench-anonymous/AptaBench_dataset.apty
APTY
Dataset from the paper "Towards Human Understanding of Paraphrase Types in ChatGPT" (https://arxiv.org/abs/2407.02302). It consists of two parts: The first part (APTYbase) contains annotated paraphrases with specific atomic paraphrase types based on the ETPC dataset. The second part (APTYranked) consists of human preferences ranking paraphrases with specific atomic paraphrase types.
The code to generate the paraphrase candidates can be found at… See the full description on the dataset page: https://huggingface.co/datasets/worta/apty.crates-dataset
CratesDataset
This revision publishes the immutable crates.io metadata snapshot 2026-07-25 generated by CratesDataset 0.1.0. It contains eight extractive or deterministic analytical tables.
Load a table
from datasets import load_dataset
versions = load_dataset(
"Aptlantis/crates-dataset", "crate_versions", split="train"
)
Parquet is the configured Hub representation used by the dataset viewer. Matching JSONL-Zstandard files are available under data/jsonl/ as… See the full description on the dataset page: https://huggingface.co/datasets/Aptlantis/crates-dataset.AptaBench_dataset
AptaBench
AptaBench is a benchmark for aptamer–small-molecule interaction prediction. It contains curated DNA/RNA aptamer–ligand pairs with standardized sequences, canonical SMILES, experimentally grounded active/inactive labels, quantitative affinity values where available, and fixed leakage-aware evaluation splits.
This repository is provided for anonymous peer review. Author identities, affiliations, acknowledgements, citation information, and non-anonymous project… See the full description on the dataset page: https://huggingface.co/datasets/swampfireee/AptaBench_dataset.cluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-sigmoid-details
Dataset Card for Evaluation run of cluebbers/Llama-3.1-8B-paraphrase-type-generation-apty-sigmoid
Dataset automatically created during the evaluation run of model cluebbers/Llama-3.1-8B-paraphrase-type-generation-apty-sigmoid
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-sigmoid-details.record-cube1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 125,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aptni27/record-cube1.cluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-sigmoidjob-aptitude
Job Role Aptitude Dataset
A dataset of aptitude and assessment questions for training and evaluating models that generate multiple-choice questions from job context: general role, skills, experience, and employment type.
Coverage includes Software Developer, Doctor, Pharmacist, Nurse, Teacher, Data Analyst, engineers, managers, and many other general professions.
Dataset Files
File
Purpose
aptitude_train.jsonl
Training split
aptitude_test.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Lumixis/job-aptitude.aptitude-test-jobresults_on_code_generation_qwen3-14B-base_float32APTO-SorryBench-JA
APTO-SorryBench-JA
APTO-SorryBench-JA is a Japanese translated and annotated version of the original Sorry-Bench dataset for AI safety evaluation research.
Sorry-Bench is a benchmark designed to evaluate whether Large Language Models (LLMs) appropriately refuse harmful requests while still providing helpful responses to safe requests.
The original English prompts are preserved alongside the Japanese translations to improve traceability and facilitate comparison with the original… See the full description on the dataset page: https://huggingface.co/datasets/APTO-001/APTO-SorryBench-JA.cluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-ipoafrica-apt-espionage
Cyber Espionage & State-Sponsored APT (Africa) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet, optimized-parquet - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-apt-espionage.cluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-ipo-details
Dataset Card for Evaluation run of cluebbers/Llama-3.1-8B-paraphrase-type-generation-apty-ipo
Dataset automatically created during the evaluation run of model cluebbers/Llama-3.1-8B-paraphrase-type-generation-apty-ipo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-ipo-details.tokenized_ds_stats_apt4
