datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
webcode2m_purifiedWebCode2M: A Real-World Dataset for Code Generation from Webpage Designs
Features:
image: the screenshot of the webpage.
bbox: the layout information, i.e., the bounding boxes (Bbox) of all the elements in the webpage, which contains the size, position, and hierarchy information.
text: the webpage code text including HTML/CSS code.
scale: the scale of the screenshot, in the format [width, height].
lang: the main language of the text content displayed on the rendered page (excluding HTML/CSS… See the full description on the dataset page: https://huggingface.co/datasets/xcodemind/webcode2m_purified.webcode2mWebCode2M: A Real-World Dataset for Code Generation from Webpage Designs with Layouts
(This dataset is also called Vision2UI.)
Automatically generating webpage code from webpage designscan significantly reduce the workload of front-end developers, andrecent Multimodal Large Language Models (MLLMs) have shownpromising potential in this area. However, our investigation revealsthat most existing MLLMs are constrained by the absence of highquality, large-scale, real-world datasets, resulting in… See the full description on the dataset page: https://huggingface.co/datasets/xcodemind/webcode2m.webcode2m-natural-promptswebcoder-250kwebcode2m-scored-promptswebcode2m_testLeroyDyer___Spydaz_Web_AI_AGI_R1_OmG_Coder-details
Dataset Card for Evaluation run of LeroyDyer/_Spydaz_Web_AI_AGI_R1_OmG_Coder
Dataset automatically created during the evaluation run of model LeroyDyer/_Spydaz_Web_AI_AGI_R1_OmG_Coder
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer___Spydaz_Web_AI_AGI_R1_OmG_Coder-details.vibe-coded-web-apps
🌀 vibe-coded-web-apps
A synthetic dataset of 20,656 web applications generated by a diverse set of frontier and open-weight language models.Each app is a fully structured project—ranging from REST/GraphQL APIs to full-stack production-grade applications—capturing the unique "vibe" and coding style of the model that created it.
🌍 Overview
The vibe-coded-web-apps dataset is a curated collection of complete web applications generated by multiple Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/harisec/vibe-coded-web-apps.Gemma4-E2B-SFT-WebCode
Gemma4-E2B-SFT-WebCode
Synthetic frontend web development dataset. Natural language component description → production-ready code.
Frameworks: React, TypeScript, Tailwind CSS, Vanilla HTML/CSS/JS.
Components: Navigation, forms, modals, data tables, charts, infinite scroll, etc.
Format: ShareGPT/ChatML. Includes accessibility attributes and comments.
Use: Fine-tune models for frontend copilot tasks.
Generator: DuoNeural/TurboGemma4E2B, temperature 0.65.
LeroyDyer___Spydaz_Web_AI_AGI_R1_Student_Coder-details
Dataset Card for Evaluation run of LeroyDyer/_Spydaz_Web_AI_AGI_R1_Student_Coder
Dataset automatically created during the evaluation run of model LeroyDyer/_Spydaz_Web_AI_AGI_R1_Student_Coder
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer___Spydaz_Web_AI_AGI_R1_Student_Coder-details.Gemma4-E2B-SFT-WebCode
Gemma4-E2B-SFT-WebCode
Synthetic frontend web development dataset. Natural language component description → production-ready code.
Frameworks: React, TypeScript, Tailwind CSS, Vanilla HTML/CSS/JS.
Components: Navigation, forms, modals, data tables, charts, infinite scroll, etc.
Format: ShareGPT/ChatML. Includes accessibility attributes and comments.
Use: Fine-tune models for frontend copilot tasks.
Generator: DuoNeural/TurboGemma4E2B, temperature 0.65.
LeroyDyer___Spydaz_Web_AI_AGI_R1_Teacher_Coder-details
Dataset Card for Evaluation run of LeroyDyer/_Spydaz_Web_AI_AGI_R1_Teacher_Coder
Dataset automatically created during the evaluation run of model LeroyDyer/_Spydaz_Web_AI_AGI_R1_Teacher_Coder
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer___Spydaz_Web_AI_AGI_R1_Teacher_Coder-details.webcode2m-scored-prompts-gpt-osswebcodecs-codec-support
The upscaler.video Codec Support Dataset
Dataset Summary
This dataset contains real-world audio and video codec support data passively collected from 1,142,586 unique user sessions at free.upscaler.video. It uses the WebCodecs API to detect both encoding and decoding support for over 1,000 codec variants across major audio and video codec families (H.264, H.265, VP8, VP9, AV1, AAC, Opus, etc.).
Total Tests: 363,330,358
Sessions: 1,142,586
Unique Codecs: 1,087
Collection… See the full description on the dataset page: https://huggingface.co/datasets/katana-video/webcodecs-codec-support.webcode2m-improved-promptsunified-webcode-datasetwebcode2m-with-reasoningweb-coder-500mbagentic_ii_agent_Qwen3_coder_prompt_web_benchagentic_ii_agent_Qwen3_coder_prompt_web_bench_verifiedagentic_dataset_qwen3_coder_webagentic_dataset_qwen3_coder_web_nodeagentic_dataset_qwen3_coder_web-sftwebcode2m-tailwind-1mwebcode2m_tailwind_100k_previewweb-coder-deepseek-v4-flashopen-web-coder
