datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
webcode2m_purifiedWebCode2M: A Real-World Dataset for Code Generation from Webpage Designs
Features:
image: the screenshot of the webpage.
bbox: the layout information, i.e., the bounding boxes (Bbox) of all the elements in the webpage, which contains the size, position, and hierarchy information.
text: the webpage code text including HTML/CSS code.
scale: the scale of the screenshot, in the format [width, height].
lang: the main language of the text content displayed on the rendered page (excluding HTML/CSS… See the full description on the dataset page: https://huggingface.co/datasets/xcodemind/webcode2m_purified.webcode2mWebCode2M: A Real-World Dataset for Code Generation from Webpage Designs with Layouts
(This dataset is also called Vision2UI.)
Automatically generating webpage code from webpage designscan significantly reduce the workload of front-end developers, andrecent Multimodal Large Language Models (MLLMs) have shownpromising potential in this area. However, our investigation revealsthat most existing MLLMs are constrained by the absence of highquality, large-scale, real-world datasets, resulting in… See the full description on the dataset page: https://huggingface.co/datasets/xcodemind/webcode2m.webcode2m-natural-promptswebcoder-250kwebcode2m_testGemma4-E2B-SFT-WebCode
Gemma4-E2B-SFT-WebCode
Synthetic frontend web development dataset. Natural language component description → production-ready code.
Frameworks: React, TypeScript, Tailwind CSS, Vanilla HTML/CSS/JS.
Components: Navigation, forms, modals, data tables, charts, infinite scroll, etc.
Format: ShareGPT/ChatML. Includes accessibility attributes and comments.
Use: Fine-tune models for frontend copilot tasks.
Generator: DuoNeural/TurboGemma4E2B, temperature 0.65.
Gemma4-E2B-SFT-WebCode
Gemma4-E2B-SFT-WebCode
Synthetic frontend web development dataset. Natural language component description → production-ready code.
Frameworks: React, TypeScript, Tailwind CSS, Vanilla HTML/CSS/JS.
Components: Navigation, forms, modals, data tables, charts, infinite scroll, etc.
Format: ShareGPT/ChatML. Includes accessibility attributes and comments.
Use: Fine-tune models for frontend copilot tasks.
Generator: DuoNeural/TurboGemma4E2B, temperature 0.65.
webcode2m-scored-prompts-gpt-osswebcodecs-codec-support
The upscaler.video Codec Support Dataset
Dataset Summary
This dataset contains real-world audio and video codec support data passively collected from 1,142,586 unique user sessions at free.upscaler.video. It uses the WebCodecs API to detect both encoding and decoding support for over 1,000 codec variants across major audio and video codec families (H.264, H.265, VP8, VP9, AV1, AAC, Opus, etc.).
Total Tests: 363,330,358
Sessions: 1,142,586
Unique Codecs: 1,087
Collection… See the full description on the dataset page: https://huggingface.co/datasets/katana-video/webcodecs-codec-support.webcode2m-improved-promptsunified-webcode-datasetwebcode2m-with-reasoningweb-coder-500mbweb-coder-deepseek-v4-flash
