datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AuroraCap-trainset
AuroraCap Trainset
Resources
Website
arXiv: Paper
GitHub: Code
Huggingface: AuroraCap Model
Huggingface: VDC Benchmark
Huggingface: Trainset
Features
We use over 20 million high-quality image/video-text pairs to train AuroraCap in three stages.
Pretraining stage. We first align visual features with the word embedding space of LLMs. To achieve this, we freeze the pretrained ViT and LLM, training solely the vision-language connector.
Vision stage. We… See the full description on the dataset page: https://huggingface.co/datasets/wchai/AuroraCap-trainset.Easy-Turn-Trainset
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
Guojian Li1, Chengyou Wang1, Hongfei Xue1,
Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2,
Yuke Lin2, Wenjie Li2, Longshuai Xiao2,
Zhonghua Fu1,╀, Lei Xie1,╀
1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University
2 Huawei Technologies, China
🎤 Demo Page
🤖 Easy Turn Model
📑 Paper
🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/Easy-Turn-Trainset.TrainSet
SIQA TrainSet
A standardized TrainSet for the Scientific Image Quality Assessment (SIQA), designed to train multimodal models on two core tasks SIQA-U & SIQA-S.
NOTICE:Due to Hugging Face Datasets' automatic caching mechanism, the local disk usage can be up to ~20× larger than the actual dataset size.
You only need to download the 'images/' folder to run inference or evaluation, please ignore the .parquet in 'SIQA-S/' and 'SIQA-U'
To Quick Start, you download by git:
git clone… See the full description on the dataset page: https://huggingface.co/datasets/SIQA/TrainSet.FakeVV_trainsetdanish-trainsConceptFormer-Trainset
ConceptFormer Trainset
Grounded visual-document retrieval training data for ConceptFormer. The release contains
37,966 training samples and an image-complete, deduplicated corpus of 12,232 positive
document images. Multi-image relations are preserved through each sample's
relevant_doc_ids list.
from datasets import load_dataset
annotations = load_dataset(
"parquet",
data_files="hf://datasets/hmhm1229/ConceptFormer-Trainset/data/train.parquet",
split="train",
)… See the full description on the dataset page: https://huggingface.co/datasets/hmhm1229/ConceptFormer-Trainset.mossling-trainsetEasy-Turn-Trainset
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
Guojian Li1, Chengyou Wang1, Hongfei Xue1,
Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2,
Yuke Lin2, Wenjie Li2, Longshuai Xiao2,
Zhonghua Fu1,╀, Lei Xie1,╀
1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University
2 Huawei Technologies, China
🎤 Demo Page
🤖 Easy Turn Model
📑 Paper
🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/0x3/Easy-Turn-Trainset.trainSnowScience-T2I-Trainset
Science-T2I Trainset
Resources
Website
arXiv: Paper
GitHub: Code
Huggingface: SciScore
Huggingface: Science-T2I-S&C Benchmark
Training Data
The data curation process involved a multi-stage approach to generate a dataset of 40,000 images, each with a resolution of 1024x1024.
Task Definition and Template Design: We began by selecting specific scientific tasks and crafting templates for three distinct prompt types: implicit, explicit, and superficial.… See the full description on the dataset page: https://huggingface.co/datasets/Jialuo21/Science-T2I-Trainset.train-shadow-blockedTrainstrain-shadows-filteredspatial-trainset-3kpb_trainsetpb_trainset-1pb_trainset-2remoteClip-trainsetlana3-trainset
