CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SEACrowd /sea-vl_crawling SEA-VL: A Multicultural Vision-Language Dataset for Southeast Asia Paper: Crowdsource, Crawl, or Generate? Creating SEA-VL, A Multicultural Vision-Language Dataset for Southeast Asia Dataset: SEA-VL Collection on HuggingFace Code: SEA-VL Experiment | SEA-VL Image Collection What is SEA-VL? Following the success of our SEACrowd project, we’re excited to announce SEA-VL, a new open-source initiative to create high-quality vision-language datasets specifically for… See the full description on the dataset page: https://huggingface.co/datasets/SEACrowd/sea-vl_crawling.image1M<n<10M5 likes473 downloads2y agoHugging Face02k-l-lambda /imslp-crawling Top Composers Composer folders size size composer 1017G / 35G /(1959)Daniel Léo Simpson 31G /(1756)Wolfgang Amadeus Mozart 24G /(1685)Johann Sebastian Bach 21G /(1770)Ludwig van Beethoven 17G /(1699)Johann Adolph Hasse 16G /(1678)Antonio Vivaldi 16G /(1670)Antonio Caldara 14G /(1685)George Frideric Handel 14G /(1683)Christoph Graupner 13G /(1975)Carlotta Ferrari 13G /(1732)Joseph Haydn 11G /(1792)Gioacchino Rossini 10G /(1948)Michel… See the full description on the dataset page: https://huggingface.co/datasets/k-l-lambda/imslp-crawling.tabular100K<n<1M1 likes121 downloads7mo agoHugging Face03projecte-aina /catalan_general_crawling Dataset Card for Catalan General Crawling Dataset Summary The Catalan General Crawling Corpus is a 435-million-token web corpus of Catalan built from the web. It has been obtained by crawling the 500 most popular .cat and .ad domains during July 2020. It consists of 434,817,705 tokens, 19,451,691 sentences and 1,016,114 documents. Documents are separated by single new lines. It is a subcorpus of the Catalan Textual Corpus. This work is licensed under a Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/catalan_general_crawling.textfill-mask100K<n<1M0 likes76 downloads2y agoHugging Face04Wissam42 /SEA-VL-Crawling-I2T SeaVLCrawlingI2TRetrieval An MTEB dataset Massive Text Embedding Benchmark SEA-VL crawling is a large-scale Southeast Asia–focused image–caption collection (~1.27M web-crawled culturally relevant pairs). For MTEB evaluation we deterministically downsample to 2048 image–caption pairs via a seeded streaming shuffle (buffer=10000), using the first non-empty caption per image. Queries are images; the corpus contains captions (image→text retrieval). Task category… See the full description on the dataset page: https://huggingface.co/datasets/Wissam42/SEA-VL-Crawling-I2T.imageother1K<n<10K0 likes37 downloads2mo agoHugging Face05Wissam42 /SEA-VL-Crawling-T2I SeaVLCrawlingT2IRetrieval An MTEB dataset Massive Text Embedding Benchmark SEA-VL crawling is a large-scale Southeast Asia–focused image–caption collection (~1.27M web-crawled culturally relevant pairs). For MTEB evaluation we deterministically downsample to 2048 image–caption pairs via a seeded streaming shuffle (buffer=10000), using the first non-empty caption per image. Queries are captions; the corpus contains images (text→image retrieval). Task category… See the full description on the dataset page: https://huggingface.co/datasets/Wissam42/SEA-VL-Crawling-T2I.imageother1K<n<10K0 likes35 downloads2mo agoHugging Face06mteb /SEA-VL-Crawling-I2T SeaVLCrawlingI2TRetrieval An MTEB dataset Massive Text Embedding Benchmark SEA-VL crawling is a large-scale Southeast Asia–focused image–caption collection (~1.27M web-crawled culturally relevant pairs). For MTEB evaluation we deterministically downsample to 2048 image–caption pairs via a seeded streaming shuffle (buffer=10000), using the first non-empty caption per image. Queries are images; the corpus contains captions (image→text retrieval). Task category… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SEA-VL-Crawling-I2T.imageother1K<n<10K0 likes34 downloads2mo agoHugging Face07SEACrowd /sea-vl-crawling-metadatatabular1M<n<10M0 likes27 downloads1y agoHugging Face08mteb /SEA-VL-Crawling-T2I SeaVLCrawlingT2IRetrieval An MTEB dataset Massive Text Embedding Benchmark SEA-VL crawling is a large-scale Southeast Asia–focused image–caption collection (~1.27M web-crawled culturally relevant pairs). For MTEB evaluation we deterministically downsample to 2048 image–caption pairs via a seeded streaming shuffle (buffer=10000), using the first non-empty caption per image. Queries are captions; the corpus contains images (text→image retrieval). Task category… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SEA-VL-Crawling-T2I.imageother1K<n<10K0 likes25 downloads2mo agoHugging Face09SEACrowd /sea-vl-crawling-metadata-non-duptabular100K<n<1M0 likes20 downloads1y agoHugging Face10SEACrowd /sea-vl-crawling-metadata-reward-filtered_3.0tabular100K<n<1M0 likes16 downloads1y agoHugging Face11SEACrowd /sea-vl-crawling-metadata-reward-filtered_4.0tabular100K<n<1M0 likes14 downloads1y agoHugging Face12SEACrowd /sea-vl-crawling-metadata-filteredtabular100K<n<1M0 likes12 downloads1y agoHugging Face13bigscience-data /roots_ca_catalan_government_crawlinggatedROOTS Subset: roots_ca_catalan_government_crawling Catalan Government Crawling Dataset uid: catalan_government_crawling Description The Catalan Government Crawling Corpus is a 39-million-token web corpus of Catalan built from the web. It has been obtained by crawling the .gencat domain and subdomains, belonging to the Catalan Government during September and October 2020. It consists of 39.117.909 tokens, 1.565.433 sentences and 71.043 documents. Documents are separated… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_ca_catalan_government_crawling.text10K<n<100K0 likes5 downloads4y agoHugging Face14xodhks /crawling-emotions-in-google-testimagen<1K0 likes4 downloads2y agoHugging Face15xodhks /crawling-emotions-in-google-trainimage1K<n<10K0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.