crowd-sourcing
sea-vl_crowdsourcing
SEA-VL: A Multicultural Vision-Language Dataset for Southeast Asia
Paper: Crowdsource, Crawl, or Generate? Creating SEA-VL, A Multicultural Vision-Language Dataset for Southeast Asia
Dataset: SEA-VL Collection on HuggingFace
Code: SEA-VL Experiment | SEA-VL Image Collection
What is SEA-VL?
Following the success of our SEACrowd project, we’re excited to announce SEA-VL, a new open-source initiative to create high-quality vision-language datasets specifically for… See the full description on the dataset page: https://huggingface.co/datasets/SEACrowd/sea-vl_crowdsourcing.SEA-VL-Crowdsourcing
Team and Homepage
Official Website: https://aienthusiasm.vn
Hugging Face Organization: https://huggingface.co/ai-enthusiasm-community
Contact
If you encounter any issues with the dataset or have any inquiries, please feel free to reach out to us via email at: aienthusiasm.team@gmail.com
Dataset Structure
The dataset is provided in a flattened tabular format, optimized for the Hugging Face Dataset Viewer and high-speed Parquet processing.
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ai-enthusiasm-community/SEA-VL-Crowdsourcing.sea-vl_crowdsourcing_vqasea-vl_crowdsourcing_translationssea-vl_crowdsourcing_id
SEA-VL Crowdsourcing — Indonesia (Bahasa Indonesia & Javanese)
This dataset is an Indonesia-language subset of the SEACrowd/sea-vl_crowdsourcing dataset, prepared and redistributed by KORIKA-AI.
It contains all crowdsourced samples whose native_lang is Indonesian (ind) or Javanese (jav):
Language
Samples
Indonesian (ind)
3,086
Javanese (jav)
10
Total
3,096
The schema, fields, and content are unchanged from the source dataset — only the language-based… See the full description on the dataset page: https://huggingface.co/datasets/KORIKA-AI/sea-vl_crowdsourcing_id.CrowdsourcingPiedmontese
Crowdsourcing Piedmontese to Test LLMs on Non-Standard Orthography
This a dataset for machine translation, topic classification and word alignment in Piedmontese.
The main features are that it is crowd sourced and that it does not assume standard orthography.
The dataset (including the raw data) is also available here: http://hdl.handle.net/11372/LRT-6086.
You can read the full paper here: https://arxiv.org/abs/2602.14675
This dataset is derived from FLORES+ and SIB-200.… See the full description on the dataset page: https://huggingface.co/datasets/ufal/CrowdsourcingPiedmontese.
