datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CrowdEvaltoxicity-multi-label-classifier
Part of a course titled "Generative AI application design & development"
https://genai.acloudfan.com/
Created from a dataset available on Kaggle.
https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/data
acl-verbatim-spans
ACL-Verbatim Span Dataset
KRLabsOrg/acl-verbatim-spans
is a dataset for query-conditioned extractive evidence selection over papers from the
ACL Anthology.
The release combines:
a gold test benchmark with manual span annotations
a larger silver training set produced from synthetic questions, retrieval, and LLM-based
span annotation
an encoder-ready config for training token-classification models directly
The underlying document collection is
KRLabsOrg/acl-anthology-md.… See the full description on the dataset page: https://huggingface.co/datasets/KRLabsOrg/acl-verbatim-spans.acl-arcvega_1_pThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state.left_arm.abs_qpos": {
"dtype": "float32",
"shape": [
7
],
"names": null,
"fps": 20
},
"observation.state.left_arm.qvel": {
"dtype": "float32",
"shape": [
7
]… See the full description on the dataset page: https://huggingface.co/datasets/acleary/vega_1_p.ACL-SRW-2025
Dataset Components
The dataset is partitioned into three discrete tables stored in CSV or Parquet format:
Questions
Recipes
Evaluation Results
Each component is described in detail below.
Questions
area
domain
question_number
An integer index uniquely identifying each question inside the knowledge domain.
translation_method
English, Google Translate, GPT-3.5-Turbo, GPT-4o, Human
question
option_a, option_b, option_c, option_d
Recipes
area… See the full description on the dataset page: https://huggingface.co/datasets/Translated-MMLU-Blind-Review/ACL-SRW-2025.ACL-ARC
Dataset Card for "ACL-ARC"
More Information needed
flare-sm-acl-longacled_event_entailment
Dataset Card for "acled_event_entailment"
More Information needed
aclacl-anthology-atomic-claims
ACL Anthology Atomic Contribution Claims (ACC)
346,010 atomic contribution claims extracted from the abstracts of 80,144 ACL Anthology papers — 423 venues, 1963–2026.
An atomic contribution claim (ACC) is a single self-contained sentence stating one concrete contribution of a paper: atomic (exactly one contribution-bearing proposition), decontextualized (pronouns resolved, meta-language removed), and falsifiable (a verifiable assertion). Unlike raw abstracts or keywords, ACCs… See the full description on the dataset page: https://huggingface.co/datasets/Hamyrappy/acl-anthology-atomic-claims.aclAnthology-9k-filtered
Description
This dataset contains filtered ACL papers between certain abstract length ranges. The data includes columns such as paper_name, year, venue, url, bibkey, and cite_acl.
We also have the filtered_dataset.jsonl that holds the main text info.
Note: Some records might be missing certain fields, especially bibkey or cite_acl. We plan to fill them via partial manual / fuzzy matching.
License
Materials prior to 2016: CC BY-NC-SA 3.0.
Materials from… See the full description on the dataset page: https://huggingface.co/datasets/yilmazzey/aclAnthology-9k-filtered.acl-voice-cloning-fr-expanded1embedded_movies_smallThis dataset was created from the HuggingFace dataset AIatMongoDB/embedded_movies
Why was it needed?
The original dataset is close to 25 GB, for learning and experiments it is an overkill
Data in the dataset needs to be cleaned up e.g., some features are Null that requires extra care
Some of the embeddings are missing
How to use?
Use for sentiment analysis
Text similarity (plot)
Embeddings : ready to use with vector DB & search libraries
dataset_info:
features:
- name:… See the full description on the dataset page: https://huggingface.co/datasets/acloudfan/embedded_movies_small.acl-anthology_sumteste_toxicidade_acl_250331flare-sm-acl-long-instructionACLing-dataprompt-sensitivity-codegen
Anonymous Prompt Sensitivity Dataset
This package contains model generations and evaluation outcomes for an anonymized
submission on prompt sensitivity in few-shot code generation.
What is included
prompt_sensitivity_dataset.jsonl: one row per generated sample
prompt_sensitivity_dataset.csv: tabular view of the same rows
prompt_sensitivity_dataset.parquet: columnar copy when parquet support is available
prompt_variant_spec.json: machine-readable description of the prompt… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-acl26/prompt-sensitivity-codegen.flare-sm-acl-longacl-crown-analysis-data
