datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
acl-anthology-md
ACL Anthology Markdown Corpus
A snapshot of the ACL Anthology consisting of bibliographic metadata for 120,034 papers and full-text markdown conversions of 114,484 papers (≈95% of the catalogue, the remainder are frontmatter, abstract-only entries, or papers without an available PDF).
This corpus is the document collection used by ACL-Verbatim, a hallucination-free question-answering system for NLP research papers built on top of VerbatimRAG.
Configurations
The… See the full description on the dataset page: https://huggingface.co/datasets/KRLabsOrg/acl-anthology-md.x-ray_acl-dataset
Data Source
https://universe.roboflow.com/roboflow-100/acl-x-ray
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact
mahadise01@gmail.com
Linkdin: https://www.linkedin.com/in/mahadise01
Github: https://github.com/Mahadih534
acl-x-ray
Dataset Card for acl-x-ray
** The original COCO dataset is stored at dataset.tar.gz**
Dataset Summary
acl-x-ray
Supported Tasks and Leaderboards
object-detection: The dataset can be used to train a model for Object Detection.
Languages
English
Dataset Structure
Data Instances
A data point comprises an image and its object annotations.
{
'image_id': 15,
'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=640x640… See the full description on the dataset page: https://huggingface.co/datasets/Francesco/acl-x-ray.qwen-acl_eval-qwen_predacl_evalgemma4-acl_eval-gemma4_predACL-fig
Dataset Card for ACLFig Dataset
Dataset Summary
1758 total labelled images
The scientific figures dataset contains 1758 scientific figures extracted from 890 research papers(ACL). The scientific figures are in png format.
The dataset has been classified into 19 categories. These are
Algorithms
Architecture/Pipeline diagrams
Bar charts
Box Plots
Confusion Matrix
Graph
Line Chart
Maps
Natural Images
Neural Networks
NLP rules/grammar
Pie chart
Scatter Plot… See the full description on the dataset page: https://huggingface.co/datasets/citeseerx/ACL-fig.aclcifar20
Dataset Card for ACLCIFAR20
This Complementary labeled CIFAR20 dataset contains auto-labeled complementary labels for all 50000 images in the training split of CIFAR20. We group 4-6 categories as a superclass and collect the complementary labels of these 20 superclasses.
For more details, please visit our github or paper.
Dataset Structure
Data Instances
A sample from the training set is provided below:
{
'images': <PIL.PngImagePlugin.PngImageFile… See the full description on the dataset page: https://huggingface.co/datasets/ntucllab/aclcifar20.aclmicroimagenet20
Dataset Card for ACLMicroImageNet10
This Complementary labeled MicroImageNet20 dataset contains 3 human-annotated complementary labels for all 10000 images in the training split of TinyImageNet200.
For more details, please visit our github or paper.
Dataset Structure
Data Instances
A sample from the training set is provided below:
{
'images': <PIL.PngImagePlugin.PngImageFile image mode=RGB size=64x64 at 0x796754E75E70>,
'ord_labels': 0… See the full description on the dataset page: https://huggingface.co/datasets/ntucllab/aclmicroimagenet20.ACLFigVQAembedded_movies_smallThis dataset was created from the HuggingFace dataset AIatMongoDB/embedded_movies
Why was it needed?
The original dataset is close to 25 GB, for learning and experiments it is an overkill
Data in the dataset needs to be cleaned up e.g., some features are Null that requires extra care
Some of the embeddings are missing
How to use?
Use for sentiment analysis
Text similarity (plot)
Embeddings : ready to use with vector DB & search libraries
dataset_info:
features:
- name:… See the full description on the dataset page: https://huggingface.co/datasets/acloudfan/embedded_movies_small.aclcifar10
Dataset Card for ACLCIFAR10
This Complementary labeled CIFAR10 dataset contains auto-labeled complementary labels for all 50000 images in the training split of CIFAR10.
For more details, please visit our github or paper.
Dataset Structure
Data Instances
A sample from the training set is provided below:
{
'images': <PIL.PngImagePlugin.PngImageFile image mode=RGB size=32x32 at 0x744F7C209DE0>,
'ord_labels': 6,
'cl_labels': [8, 1, 9, 9, 0, 9, 0, 2… See the full description on the dataset page: https://huggingface.co/datasets/ntucllab/aclcifar10.acl_eval_docvqaintern-acl_eval-internvl_predacl_dataacl_predmmscibench-anonymous
Dataset Name
Dataset Summary
MMSciBench dataset focuses on mathematics and physics that evaluates scientific reasoning capabilities.
Dataset Structure
Data Instances
QA_has_img_with_categories.csv: Contains data of Q&A questions (with images) of the subject indicated by the folder name.
QA_no_img_with_categories.csv: Contains data of text-only Q&A questions of the subject indicated by the folder name.
multi_choice_has_img_with_categories.csv:… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-acl-submission/mmscibench-anonymous.aclmicroimagenet10
Dataset Card for ACLMicroImageNet10
This Complementary labeled MicroImageNet10 dataset contains 3 human-annotated complementary labels for all 5000 images in the training split of TinyImageNet200.
For more details, please visit our github or paper.
Dataset Structure
Data Instances
A sample from the training set is provided below:
{
'images': <PIL.PngImagePlugin.PngImageFile image mode=RGB size=64x64 at 0x77CB6DF65F30>,
'ord_labels': 0… See the full description on the dataset page: https://huggingface.co/datasets/ntucllab/aclmicroimagenet10.waste-wizard-materials-listAncientVision-3T
AncientVision-3T: A Hierarchical Benchmark for Ancient Chinese Document Vision-Language Tasks
Dataset Summary
AncientVision-3T is a benchmark designed to systematically evaluate the cognitive capabilities of Vision Language Models (VLMs) in the context of Ancient Chinese Documents. Unlike general benchmarks, AncientVision-3T employs a hierarchical task design to decouple and analyze model capabilities based on cognitive complexity.
Processing ancient Chinese documents… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-acl-2026/AncientVision-3T.
