CoolFace
18 results

crawling

SEACrowd /sea-vl_crawling SEA-VL: A Multicultural Vision-Language Dataset for Southeast Asia Paper: Crowdsource, Crawl, or Generate? Creating SEA-VL, A Multicultural Vision-Language Dataset for Southeast Asia Dataset: SEA-VL Collection on HuggingFace Code: SEA-VL Experiment | SEA-VL Image Collection What is SEA-VL? Following the success of our SEACrowd project, we’re excited to announce SEA-VL, a new open-source initiative to create high-quality vision-language datasets specifically for… See the full description on the dataset page: https://huggingface.co/datasets/SEACrowd/sea-vl_crawling.image1M<n<10M5 likes474 downloads2y agoHugging Facek-l-lambda /imslp-crawling Top Composers Composer folders size size composer 1017G / 35G /(1959)Daniel Léo Simpson 31G /(1756)Wolfgang Amadeus Mozart 24G /(1685)Johann Sebastian Bach 21G /(1770)Ludwig van Beethoven 17G /(1699)Johann Adolph Hasse 16G /(1678)Antonio Vivaldi 16G /(1670)Antonio Caldara 14G /(1685)George Frideric Handel 14G /(1683)Christoph Graupner 13G /(1975)Carlotta Ferrari 13G /(1732)Joseph Haydn 11G /(1792)Gioacchino Rossini 10G /(1948)Michel… See the full description on the dataset page: https://huggingface.co/datasets/k-l-lambda/imslp-crawling.tabular100K<n<1M1 likes134 downloads7mo agoHugging Faceprojecte-aina /catalan_general_crawling Dataset Card for Catalan General Crawling Dataset Summary The Catalan General Crawling Corpus is a 435-million-token web corpus of Catalan built from the web. It has been obtained by crawling the 500 most popular .cat and .ad domains during July 2020. It consists of 434,817,705 tokens, 19,451,691 sentences and 1,016,114 documents. Documents are separated by single new lines. It is a subcorpus of the Catalan Textual Corpus. This work is licensed under a Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/catalan_general_crawling.textfill-mask100K<n<1M0 likes68 downloads2y agoHugging Facealvinrifky /Crawling-MKN_10 likes60 downloads4mo agoHugging Faceprojecte-aina /catalan_government_crawling Dataset Card for Catalan Government Crawling Dataset Summary The Catalan Government Crawling Corpus is a 39-million-token web corpus of Catalan built from the web. It has been obtained by crawling the .gencat domain and subdomains, belonging to the Catalan Government during September and October 2020. It consists of 39,117,909 tokens, 1,565,433 sentences and 71,043 documents. Documents are separated by single new lines. It is a subcorpus of the Catalan Textual Corpus.… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/catalan_government_crawling.textfill-mask10K<n<100K1 likes57 downloads2y agoHugging Facemteb /SEA-VL-Crawling-T2I SeaVLCrawlingT2IRetrieval An MTEB dataset Massive Text Embedding Benchmark SEA-VL crawling is a large-scale Southeast Asia–focused image–caption collection (~1.27M web-crawled culturally relevant pairs). For MTEB evaluation we deterministically downsample to 2048 image–caption pairs via a seeded streaming shuffle (buffer=10000), using the first non-empty caption per image. Queries are captions; the corpus contains images (text→image retrieval). Task category… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SEA-VL-Crawling-T2I.imageother1K<n<10K0 likes52 downloads2mo agoHugging Face