CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01IAmFuch /viet-cultural-vqa 🇻🇳 Vietnamese Cultural VQA Dataset 📖 Dataset Description The Vietnamese Cultural VQA Dataset is a comprehensive multimodal dataset designed for Visual Question Answering (VQA) tasks focused on Vietnamese cultural heritage. This dataset aims to bridge the gap in understanding and preserving Vietnamese culture through AI-powered visual understanding and question answering. 🎯 Dataset Summary 📊 Total Images: 28,505 high-quality cultural images 💬 Total… See the full description on the dataset page: https://huggingface.co/datasets/IAmFuch/viet-cultural-vqa.imagevisual-question-answering10K<n<100K0 likes502 downloads5mo agoHugging Face02iamfadi /de-multi-legaltext100K<n<1M0 likes146 downloads2y agoHugging Face03iamfortytwo /MusicAVQA-A2V-Retrieval MusicAVQA-A2V-Retrieval This is a derived retrieval benchmark from the test split of mteb/MUSIC-AVQA_cls-preprocessed at revision 29f50ae80ad4e8c1cfdbc0148aefe6fe050833dd. It uses audio queries and video corpus items. Construction The source clips are labelled with 22 musical-instrument classes. For every class, a deterministic seed (42) selects five clips as queries and ten distinct clips as corpus items. Relevance is class membership, so each query has ten… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/MusicAVQA-A2V-Retrieval.audio1K<n<10K0 likes69 downloads12d agoHugging Face04iamfadi /us-mergedtext100K<n<1M0 likes66 downloads2y agoHugging Face05iamfortytwo /YouCook2-I2Vimagen<1K0 likes64 downloads17d agoHugging Face06iamfortytwo /fiqa-decontaminated fiqa-decontaminated (MTEB layout) Repackaging of lightonai/fiqa-decontaminated into the layout expected by MTEB: a default config holding the qrels (one split per evaluation split), alongside corpus and queries configs. Rows and columns are copied verbatim from the source dataset; only the config/split packaging differs. All credit for the decontaminated data belongs to LightOn AI, and to the original BEIR authors for the underlying benchmark. Source:… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/fiqa-decontaminated.tabulartext-retrieval10K<n<100K0 likes58 downloads14d agoHugging Face07iamfortytwo /nfcorpus-decontaminated nfcorpus-decontaminated (MTEB layout) Repackaging of lightonai/nfcorpus-decontaminated into the layout expected by MTEB: a default config holding the qrels (one split per evaluation split), alongside corpus and queries configs. Rows and columns are copied verbatim from the source dataset; only the config/split packaging differs. All credit for the decontaminated data belongs to LightOn AI, and to the original BEIR authors for the underlying benchmark. Source:… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/nfcorpus-decontaminated.texttext-retrieval10K<n<100K0 likes56 downloads14d agoHugging Face08iamfortytwo /scifact-decontaminated scifact-decontaminated (MTEB layout) Repackaging of lightonai/scifact-decontaminated into the layout expected by MTEB: a default config holding the qrels (one split per evaluation split), alongside corpus and queries configs. Rows and columns are copied verbatim from the source dataset; only the config/split packaging differs. All credit for the decontaminated data belongs to LightOn AI, and to the original BEIR authors for the underlying benchmark. Source:… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/scifact-decontaminated.tabulartext-retrieval1K<n<10K0 likes55 downloads14d agoHugging Face09iamfortytwo /YouCook2-V2Iimagen<1K0 likes52 downloads17d agoHugging Face10iamfortytwo /arguana-decontaminated arguana-decontaminated (MTEB layout) Repackaging of lightonai/arguana-decontaminated into the layout expected by MTEB: a default config holding the qrels (one split per evaluation split), alongside corpus and queries configs. Rows and columns are copied verbatim from the source dataset; only the config/split packaging differs. All credit for the decontaminated data belongs to LightOn AI, and to the original BEIR authors for the underlying benchmark. Source:… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/arguana-decontaminated.texttext-retrieval10K<n<100K0 likes52 downloads14d agoHugging Face11iamfortytwo /webis-touche2020-decontaminated webis-touche2020-decontaminated (MTEB layout) Repackaging of lightonai/webis-touche2020-decontaminated into the layout expected by MTEB: a default config holding the qrels (one split per evaluation split), alongside corpus and queries configs. Rows and columns are copied verbatim from the source dataset; only the config/split packaging differs. All credit for the decontaminated data belongs to LightOn AI, and to the original BEIR authors for the underlying benchmark. Source:… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/webis-touche2020-decontaminated.tabulartext-retrieval100K<n<1M0 likes52 downloads14d agoHugging Face12iamfortytwo /MetaWorld-MT50-I2Vimage1K<n<10K0 likes50 downloads17d agoHugging Face13iamfortytwo /scidocs-decontaminated scidocs-decontaminated (MTEB layout) Repackaging of lightonai/scidocs-decontaminated into the layout expected by MTEB: a default config holding the qrels (one split per evaluation split), alongside corpus and queries configs. Rows and columns are copied verbatim from the source dataset; only the config/split packaging differs. All credit for the decontaminated data belongs to LightOn AI, and to the original BEIR authors for the underlying benchmark. Source:… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/scidocs-decontaminated.texttext-retrieval1K<n<10K0 likes50 downloads14d agoHugging Face14iamfortytwo /quora-decontaminated quora-decontaminated (MTEB layout) Repackaging of lightonai/quora-decontaminated into the layout expected by MTEB: a default config holding the qrels (one split per evaluation split), alongside corpus and queries configs. Rows and columns are copied verbatim from the source dataset; only the config/split packaging differs. All credit for the decontaminated data belongs to LightOn AI, and to the original BEIR authors for the underlying benchmark. Source:… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/quora-decontaminated.tabulartext-retrieval100K<n<1M0 likes48 downloads14d agoHugging Face15iamfortytwo /MusicAVQA-V2A-Retrieval MusicAVQA-V2A-Retrieval This is a derived retrieval benchmark from the test split of mteb/MUSIC-AVQA_cls-preprocessed at revision 29f50ae80ad4e8c1cfdbc0148aefe6fe050833dd. It uses video queries and audio corpus items. Construction The source clips are labelled with 22 musical-instrument classes. For every class, a deterministic seed (42) selects five clips as queries and ten distinct clips as corpus items. Relevance is class membership, so each query has ten… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/MusicAVQA-V2A-Retrieval.audio1K<n<10K0 likes47 downloads12d agoHugging Face16iamfortytwo /MetaWorld-MT50-V2Iimage1K<n<10K0 likes46 downloads17d agoHugging Face17iamfortytwo /trec-covid-decontaminated trec-covid-decontaminated (MTEB layout) Repackaging of lightonai/trec-covid-decontaminated into the layout expected by MTEB: a default config holding the qrels (one split per evaluation split), alongside corpus and queries configs. Rows and columns are copied verbatim from the source dataset; only the config/split packaging differs. All credit for the decontaminated data belongs to LightOn AI, and to the original BEIR authors for the underlying benchmark. Source:… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/trec-covid-decontaminated.tabulartext-retrieval100K<n<1M0 likes45 downloads14d agoHugging Face18iamfadi /eu-mergedtext100K<n<1M0 likes39 downloads2y agoHugging Face19iamfadi /multi_eurlex_en_processedtext100K<n<1M0 likes36 downloads2y agoHugging Face20iamfadi /us-evaltext100K<n<1M0 likes18 downloads2y agoHugging Face21iamfadi /edgar_alltext10K<n<100K0 likes15 downloads2y agoHugging Face22iamfadi /edgar_all4text10K<n<100K0 likes8 downloads2y agoHugging Face23iamfadi /indian_nertext10K<n<100K0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.