datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CrowdEvalacl-anthology-md
ACL Anthology Markdown Corpus
A snapshot of the ACL Anthology consisting of bibliographic metadata for 120,034 papers and full-text markdown conversions of 114,484 papers (≈95% of the catalogue, the remainder are frontmatter, abstract-only entries, or papers without an available PDF).
This corpus is the document collection used by ACL-Verbatim, a hallucination-free question-answering system for NLP research papers built on top of VerbatimRAG.
Configurations
The… See the full description on the dataset page: https://huggingface.co/datasets/KRLabsOrg/acl-anthology-md.x-ray_acl-dataset
Data Source
https://universe.roboflow.com/roboflow-100/acl-x-ray
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact
mahadise01@gmail.com
Linkdin: https://www.linkedin.com/in/mahadise01
Github: https://github.com/Mahadih534
aclImdbacl-voice-cloning-fr-expandedtry
ACL Voice Cloning FR — Expanded Pairs
Every (reference, target) segment pair within each speaker.
Audio is embedded directly — playable on the dataset viewer.
Column
Type
Role
ref_en_voice
🎤 Audio
English audio of reference segment (voice to clone)
ref_fr_voice
🎤 Audio
French cloned audio of reference segment
ref_en_text
string
English text of reference
ref_fr_text
string
French text of reference
trg_en_voice
🎤 Audio
English audio of target segment… See the full description on the dataset page: https://huggingface.co/datasets/amanuelbyte/acl-voice-cloning-fr-expandedtry.acl-voice-cloning-fr-expanded
ACL Voice Cloning FR — Expanded Pairs
Every (reference, target) segment pair within each speaker.
Audio is embedded directly — playable on the dataset viewer.
Column
Type
Role
ref_en_voice
🎤 Audio
English audio of reference segment (voice to clone)
ref_fr_voice
🎤 Audio
French cloned audio of reference segment
ref_en_text
string
English text of reference
ref_fr_text
string
French text of reference
trg_en_voice
🎤 Audio
English audio of target segment… See the full description on the dataset page: https://huggingface.co/datasets/amanuelbyte/acl-voice-cloning-fr-expanded.flare-sm-acl
Dataset Card for "flare-sm-acl"
More Information needed
toxicity-multi-label-classifier
Part of a course titled "Generative AI application design & development"
https://genai.acloudfan.com/
Created from a dataset available on Kaggle.
https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/data
acl-verbatim-spans
ACL-Verbatim Span Dataset
KRLabsOrg/acl-verbatim-spans
is a dataset for query-conditioned extractive evidence selection over papers from the
ACL Anthology.
The release combines:
a gold test benchmark with manual span annotations
a larger silver training set produced from synthetic questions, retrieval, and LLM-based
span annotation
an encoder-ready config for training token-classification models directly
The underlying document collection is
KRLabsOrg/acl-anthology-md.… See the full description on the dataset page: https://huggingface.co/datasets/KRLabsOrg/acl-verbatim-spans.aclsum
ACLSum: A New Dataset for Aspect-based Summarization of Scientific Publications
This repository contains data for our paper "ACLSum: A New Dataset for Aspect-based Summarization of Scientific Publications" and a small
utility class to work with it.
HuggingFace datasets
You can also use Huggin Face datasets to load ACLSum (dataset link).
This would be convenient if you want to train transformer models using our dataset.
Just do,
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/sobamchan/aclsum.ACL-Accepted-Papersacl-6060
ACL 60/60
Dataset details
ACL 60/60 evaluation sets for multilingual translation of ACL 2022 technical presentations into 10 target languages.
Citation
@inproceedings{salesky-etal-2023-evaluating,
title = "Evaluating Multilingual Speech Translation under Realistic Conditions with Resegmentation and Terminology",
author = "Salesky, Elizabeth and
Darwish, Kareem and
Al-Badrashiny, Mohamed and
Diab, Mona and
Niehues, Jan"… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/acl-6060.acled-token-summaryACLED Dataset for Summarization Task - CSE635 (University at Buffalo)
Actor Description
0: N/A
1: State Forces
2: Rebel Groups
3: Political Militias
4: Identity Militias
5: Rioters
6: Protesters
7: Civilians
8: External/Other Forces
ACL-Bench-Task-Transfer
ACL Bench
ACL Bench is a reasoning benchmark developed to evaluate
structural task transfer after rule-constrained fine-tuning
of large language models.
It abstracts structural operations from Dou Dizhu into
three reasoning categories:
Hierarchy Logic
Pattern Recognition
Counter Logic
Dataset Size
300 multiple-choice questions:
100 Hierarchy Logic
100 Pattern Recognition
100 Counter Logic
Project
Paper:
Rule-Constrained Fine-Tuning and Task… See the full description on the dataset page: https://huggingface.co/datasets/farshelina/ACL-Bench-Task-Transfer.CocoScisum_ACLflare-sm-acl
Dataset Card for "flare-sm-acl"
More Information needed
acl-paper
ACL Entire
ACL Entire is a comprehensive dataset containing all papers from both ACL and Non-ACL events listed on the ACL Anthology website. This dataset includes complete bibliographic information for all years.
Features
Events Covered: Papers from ACL and Non-ACL events.
Bibliography: Includes complete bibliographic details for every paper.
Years Covered: Comprehensive data spanning all available years.
Source
All data has been compiled from the ACL… See the full description on the dataset page: https://huggingface.co/datasets/sleeping-ai/acl-paper.acl-voice-cloning-fr-cleaned-v2
ACL Voice Cloning FR — Cleaned & Expanded (V2)
Source filtered with GPU-accelerated quality checks (SNR≥10.0dB, silence≤65%),
then expanded into all (reference, target) pairs per speaker.
Filter
Threshold
Duration
1.0-20.0s
SNR
≥ 10.0 dB
Silence
≤ 65%
Text length
5-500 chars
Split
Source kept
Expanded pairs
Shards
train
748
70,006
43
test
119
7,022
7
qwen-acl_eval-qwen_predacl-6060-openvoiceACL2024ACL2025acl-arcacl_evalAgenticLM-Instruct-70kAgenticLM-Instruct-70k is a dataset created for training small agentic LLMs, using a tool schema (write_file, read_file, append_file, list_directory, delete_file, run_python, run_shell, install_package, run_tests, search_web, fetch_url)
Primarily for training coding agents, but training with this dataset should also work for any agents that will do work other than coding.
REALEC_GEC_dataset_ACL_testACL2026ACL_clearwikismallPSR_ACL2520 Messages picked from twitch, 7 streamers total.
Link to paper: https://aclanthology.org/2026.acl-short.51/
Cite (ACL): Mohammadsadegh Abolhasani, Reza Mousavi, and Paul Jen-Hwa Hu. 2026. CaBSALLM: Efficient Context-Aware Batch Annotation of Conversational Streams with Large Language Models. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 615–636, San Diego, California, United States. Association for Computational… See the full description on the dataset page: https://huggingface.co/datasets/abolhasani/PSR_ACL.
