datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
viet-cultural-vqa
🇻🇳 Vietnamese Cultural VQA Dataset
📖 Dataset Description
The Vietnamese Cultural VQA Dataset is a comprehensive multimodal dataset designed for Visual Question Answering (VQA) tasks focused on Vietnamese cultural heritage. This dataset aims to bridge the gap in understanding and preserving Vietnamese culture through AI-powered visual understanding and question answering.
🎯 Dataset Summary
📊 Total Images: 28,505 high-quality cultural images
💬 Total… See the full description on the dataset page: https://huggingface.co/datasets/IAmFuch/viet-cultural-vqa.german-news-intelligencede-multi-legalMusicAVQA-A2V-Retrieval
MusicAVQA-A2V-Retrieval
This is a derived retrieval benchmark from the test split of
mteb/MUSIC-AVQA_cls-preprocessed at
revision 29f50ae80ad4e8c1cfdbc0148aefe6fe050833dd. It uses audio queries and video corpus items.
Construction
The source clips are labelled with 22 musical-instrument classes. For every
class, a deterministic seed (42) selects five clips as queries and ten distinct
clips as corpus items. Relevance is class membership, so each query has ten… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/MusicAVQA-A2V-Retrieval.us-mergedYouCook2-I2Vfiqa-decontaminated
fiqa-decontaminated (MTEB layout)
Repackaging of lightonai/fiqa-decontaminated
into the layout expected by MTEB:
a default config holding the qrels (one split per evaluation split), alongside
corpus and queries configs.
Rows and columns are copied verbatim from the source dataset; only the
config/split packaging differs. All credit for the decontaminated data belongs to
LightOn AI, and to the original BEIR
authors for the underlying benchmark.
Source:… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/fiqa-decontaminated.nfcorpus-decontaminated
nfcorpus-decontaminated (MTEB layout)
Repackaging of lightonai/nfcorpus-decontaminated
into the layout expected by MTEB:
a default config holding the qrels (one split per evaluation split), alongside
corpus and queries configs.
Rows and columns are copied verbatim from the source dataset; only the
config/split packaging differs. All credit for the decontaminated data belongs to
LightOn AI, and to the original BEIR
authors for the underlying benchmark.
Source:… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/nfcorpus-decontaminated.scifact-decontaminated
scifact-decontaminated (MTEB layout)
Repackaging of lightonai/scifact-decontaminated
into the layout expected by MTEB:
a default config holding the qrels (one split per evaluation split), alongside
corpus and queries configs.
Rows and columns are copied verbatim from the source dataset; only the
config/split packaging differs. All credit for the decontaminated data belongs to
LightOn AI, and to the original BEIR
authors for the underlying benchmark.
Source:… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/scifact-decontaminated.YouCook2-V2Iarguana-decontaminated
arguana-decontaminated (MTEB layout)
Repackaging of lightonai/arguana-decontaminated
into the layout expected by MTEB:
a default config holding the qrels (one split per evaluation split), alongside
corpus and queries configs.
Rows and columns are copied verbatim from the source dataset; only the
config/split packaging differs. All credit for the decontaminated data belongs to
LightOn AI, and to the original BEIR
authors for the underlying benchmark.
Source:… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/arguana-decontaminated.webis-touche2020-decontaminated
webis-touche2020-decontaminated (MTEB layout)
Repackaging of lightonai/webis-touche2020-decontaminated
into the layout expected by MTEB:
a default config holding the qrels (one split per evaluation split), alongside
corpus and queries configs.
Rows and columns are copied verbatim from the source dataset; only the
config/split packaging differs. All credit for the decontaminated data belongs to
LightOn AI, and to the original BEIR
authors for the underlying benchmark.
Source:… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/webis-touche2020-decontaminated.MetaWorld-MT50-I2Vscidocs-decontaminated
scidocs-decontaminated (MTEB layout)
Repackaging of lightonai/scidocs-decontaminated
into the layout expected by MTEB:
a default config holding the qrels (one split per evaluation split), alongside
corpus and queries configs.
Rows and columns are copied verbatim from the source dataset; only the
config/split packaging differs. All credit for the decontaminated data belongs to
LightOn AI, and to the original BEIR
authors for the underlying benchmark.
Source:… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/scidocs-decontaminated.quora-decontaminated
quora-decontaminated (MTEB layout)
Repackaging of lightonai/quora-decontaminated
into the layout expected by MTEB:
a default config holding the qrels (one split per evaluation split), alongside
corpus and queries configs.
Rows and columns are copied verbatim from the source dataset; only the
config/split packaging differs. All credit for the decontaminated data belongs to
LightOn AI, and to the original BEIR
authors for the underlying benchmark.
Source:… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/quora-decontaminated.MusicAVQA-V2A-Retrieval
MusicAVQA-V2A-Retrieval
This is a derived retrieval benchmark from the test split of
mteb/MUSIC-AVQA_cls-preprocessed at
revision 29f50ae80ad4e8c1cfdbc0148aefe6fe050833dd. It uses video queries and audio corpus items.
Construction
The source clips are labelled with 22 musical-instrument classes. For every
class, a deterministic seed (42) selects five clips as queries and ten distinct
clips as corpus items. Relevance is class membership, so each query has ten… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/MusicAVQA-V2A-Retrieval.MetaWorld-MT50-V2Itrec-covid-decontaminated
trec-covid-decontaminated (MTEB layout)
Repackaging of lightonai/trec-covid-decontaminated
into the layout expected by MTEB:
a default config holding the qrels (one split per evaluation split), alongside
corpus and queries configs.
Rows and columns are copied verbatim from the source dataset; only the
config/split packaging differs. All credit for the decontaminated data belongs to
LightOn AI, and to the original BEIR
authors for the underlying benchmark.
Source:… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/trec-covid-decontaminated.eu-mergedmulti_eurlex_en_processedus-train-tokenizedus-evaledgar_allvietnamese-legal-documentiam-formedgar_all4IAM_full_imageindian_nerzsz-0.09k
