datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sciclaimeval-shared-task
SciClaimEval Shared Task: All information is available at sciclaimeval.github.io
Evaluation scripts & examples: github.com/SciClaimEval/sciclaimeval-shared-task
More Information: paper
Version Info
Please use the latest version, v1.1.
Changes from v1.0 to v1.1
Compared with v1.0, v1.1 includes the following changes.
Removed Samples
The following 20 samples have been removed:
val_tab_1594
val_tab_0067… See the full description on the dataset page: https://huggingface.co/datasets/alabnii/sciclaimeval-shared-task.Organic-Chemistry-VLM-PostTraining
Dataset Card for "Chemistry_text_to_image"
More Information needed
uber_text-Vision-QA
Dataset Card for "uber_text_qa"
More Information needed
Sujet-Vision-QA
Dataset Description 📊🔍
The Sujet-Finance-QA-Vision-100k is a comprehensive dataset containing over 100,000 question-answer pairs derived from more than 9,800 financial document images. This dataset is designed to support research and development in the field of financial document analysis and visual question answering.
Key Features:
🖼️ 9,801 unique financial document images
❓ 107,050 question-answer pairs
🇬🇧 English language
📄 Diverse financial document types… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/Sujet-Vision-QA.sciclaimeval-shared-task
SciClaimEval Shared Task: All information is available at sciclaimeval.github.io
Evaluation scripts & examples: github.com/SciClaimEval/sciclaimeval-shared-task
More Information: paper
Note on License Information
The dataset is licensed under CC BY 4.0; however, individual samples may have their own licenses.
sciclaimeval-shared-task-test
Grant Access: This dataset is available only to NTCIR participants.
SciClaimEval Shared Task: All information is available at sciclaimeval.github.io
Version Info
Please use the latest version, v1.1.
Changes from v1.0 to v1.1
Compared with v1.0, v1.1 removes the following 11 samples:
test_fig_0119
test_fig_0421
test_tab_0189
test_tab_0240
test_tab_0256
test_fig_0172
test_fig_0173
test_fig_0430
test_tab_0137
test_fig_0111
test_tab_0052… See the full description on the dataset page: https://huggingface.co/datasets/alabnii/sciclaimeval-shared-task-test.rrg24-shared-task-bionlp
✏️ Citation
@inproceedings{xu-etal-2024-overview,
title = "Overview of the First Shared Task on Clinical Text Generation: {RRG}24 and {\textquotedblleft}Discharge Me!{\textquotedblright}",
author = "Xu, Justin and
Chen, Zhihong and
Johnston, Andrew and
Blankemeier, Louis and
Varma, Maya and
Hom, Jason and
Collins, William J. and
Modi, Ankit and
Lloyd, Robert and
Hopkins, Benjamin and
Langlotz, Curtis and… See the full description on the dataset page: https://huggingface.co/datasets/StanfordAIMI/rrg24-shared-task-bionlp.EMNLP-NLLP-CODEswitch to private once done with paper contents, add the files along with rest on github :)
