CoolFace
17 results

WebWorld

Qwen /WebWorldData WebWorldData 🌐 Overview WebWorldData is a large-scale dataset of 1.06M web interaction trajectories collected from the open web, designed for training browser world models. It is the training data behind the WebWorld model series. Each trajectory consists of sequences of (state, action, next_state) transitions, where states are represented as A11y Trees extracted from real websites using Playwright. Dataset Statistics Total… See the full description on the dataset page: https://huggingface.co/datasets/Qwen/WebWorldData.texttext-generation100K<n<1M79 likes422 downloads5mo agoHugging FaceMultilingual-Multimodal-NLP /WebWorld-SFT-10k WebWorld-SFT-10k The Browser as a World Model for Self-Improving Web Code — a 10,000-example SFT release from the WebWorld training set. Each example is a single-turn improvement of an HTML artifact that was accepted by the browser-issued certificate: the browser re-executed the candidate, target progress held, and every previously verified capability was preserved. Rejected trajectories are not in the export. Why WebWorld VLM-driven self-improvement of web code… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Multimodal-NLP/WebWorld-SFT-10k.text-generation10K<n<100K0 likes123 downloads22d agoHugging FaceBigbarry /WebWorldData_copy WebWorldData 🌐 Overview WebWorldData is a large-scale dataset of 1.06M web interaction trajectories collected from the open web, designed for training browser world models. It is the training data behind the WebWorld model series. Each trajectory consists of sequences of (state, action, next_state) transitions, where states are represented as A11y Trees extracted from real websites using Playwright. Dataset Statistics Total… See the full description on the dataset page: https://huggingface.co/datasets/Bigbarry/WebWorldData_copy.texttext-generation100K<n<1M0 likes26 downloads4mo agoHugging Facealucent /mirror-WebWorldDatagated WebWorldData 🌐 Overview WebWorldData is a large-scale dataset of 1.06M web interaction trajectories collected from the open web, designed for training browser world models. It is the training data behind the WebWorld model series. Each trajectory consists of sequences of (state, action, next_state) transitions, where states are represented as A11y Trees extracted from real websites using Playwright. Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-WebWorldData.texttext-generation100K<n<1M0 likes6 downloads2mo agoHugging Face