crawling
Datasets
All datasets matching “crawling”sea-vl_crawling
SEA-VL: A Multicultural Vision-Language Dataset for Southeast Asia
Paper: Crowdsource, Crawl, or Generate? Creating SEA-VL, A Multicultural Vision-Language Dataset for Southeast Asia
Dataset: SEA-VL Collection on HuggingFace
Code: SEA-VL Experiment | SEA-VL Image Collection
What is SEA-VL?
Following the success of our SEACrowd project, we’re excited to announce SEA-VL, a new open-source initiative to create high-quality vision-language datasets specifically for… See the full description on the dataset page: https://huggingface.co/datasets/SEACrowd/sea-vl_crawling.imslp-crawling
Top Composers
Composer folders size
size
composer
1017G
/
35G
/(1959)Daniel Léo Simpson
31G
/(1756)Wolfgang Amadeus Mozart
24G
/(1685)Johann Sebastian Bach
21G
/(1770)Ludwig van Beethoven
17G
/(1699)Johann Adolph Hasse
16G
/(1678)Antonio Vivaldi
16G
/(1670)Antonio Caldara
14G
/(1685)George Frideric Handel
14G
/(1683)Christoph Graupner
13G
/(1975)Carlotta Ferrari
13G
/(1732)Joseph Haydn
11G
/(1792)Gioacchino Rossini
10G
/(1948)Michel… See the full description on the dataset page: https://huggingface.co/datasets/k-l-lambda/imslp-crawling.catalan_general_crawling
Dataset Card for Catalan General Crawling
Dataset Summary
The Catalan General Crawling Corpus is a 435-million-token web corpus of Catalan built from the web. It has been obtained by crawling the 500 most popular .cat and .ad domains during July 2020. It consists of 434,817,705 tokens, 19,451,691 sentences and 1,016,114 documents. Documents are separated by single new lines. It is a subcorpus of the Catalan Textual Corpus.
This work is licensed under a Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/catalan_general_crawling.Crawling-MKN_1catalan_government_crawling
Dataset Card for Catalan Government Crawling
Dataset Summary
The Catalan Government Crawling Corpus is a 39-million-token web corpus of Catalan built from the web. It has been obtained by crawling the .gencat domain and subdomains, belonging to the Catalan Government during September and October 2020. It consists of 39,117,909 tokens, 1,565,433 sentences and 71,043 documents. Documents are separated by single new lines. It is a subcorpus of the Catalan Textual Corpus.… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/catalan_government_crawling.SEA-VL-Crawling-T2I
SeaVLCrawlingT2IRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
SEA-VL crawling is a large-scale Southeast Asia–focused image–caption collection (~1.27M web-crawled culturally relevant pairs). For MTEB evaluation we deterministically downsample to 2048 image–caption pairs via a seeded streaming shuffle (buffer=10000), using the first non-empty caption per image. Queries are captions; the corpus contains images (text→image retrieval).
Task category… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SEA-VL-Crawling-T2I.
