datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Web_Scraper_Datarephrased-web-data-quality-study
Rephrased Web Data Quality Study
LLM-as-judge evaluation of ~4,000 examples from HuggingFaceFW/finephrase (1,000 sampled per split, 86 dropped due to judge parse failures, 3,914 successfully evaluated).
Judge: Claude Sonnet 4.6 via OpenRouter | Cost: ~$45
Quality Scores (1-5 scale)
Metric
FAQ (n=965)
Table (n=979)
Tutorial (n=976)
Math (n=994)
Faithfulness
1.82
1.72
1.90
1.49
Info preservation
1.93
1.64
1.99
1.47
Appropriateness
3.54
2.87
2.48
1.67… See the full description on the dataset page: https://huggingface.co/datasets/ratishsp/rephrased-web-data-quality-study.scraped-web-dataweb-search-datasetWeb_FileStructure_DataSet_100k
Dataset Name:
Web File Structure Dataset
Description:
This dataset is designed to train AI models on best practices for organizing files in web development projects. It includes 100,000 examples that cover the structure and conventions of HTML, CSS, JavaScript, and other web-related files. Each example consists of a prompt and a corresponding completion, providing comprehensive guidance on how to organize web project files effectively.
Key Features:… See the full description on the dataset page: https://huggingface.co/datasets/Juliankrg/Web_FileStructure_DataSet_100k.lexi-coding-web-clean-datasets
lexi-coding-web-clean-datasets
lexi-coding-web-clean-datasets by Reallexi LLC AI Model Builder — llm.reallexi.io
Copyright (c) 2026 Reallexi LLC. All rights reserved.
A retrieval index built by Reallexi AI Model Builder: source text chunked, embedded, and stored for nearest-neighbor retrieval. This is not a causal-language-model checkpoint and cannot be loaded with AutoModelForCausalLM.
Contents
Source data… See the full description on the dataset page: https://huggingface.co/datasets/reallexi/lexi-coding-web-clean-datasets.O2-QA-Web_dataweb_a11y_datasetWeb-dev-large-dataset
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [Cynnix]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper… See the full description on the dataset page: https://huggingface.co/datasets/cynnix69/Web-dev-large-dataset.bop-webdataset-shardsweb-builder-dataset
Web Builder Dataset — Fine-tuning untuk LLM
Dataset untuk fine-tune model pada pembangunan web builder (WeWeb clone) menggunakan Nuxt 4 + Tailwind CSS 4 + Cloudflare.
📊 Dataset Overview
Metric
Value
Total examples
500
Train
400 (80%)
Eval
100 (20%)
Format
ChatML JSONL
Language
UI: Bahasa Malaysia, Code: English
🎯 Fokus Dataset
Kategori
Bilangan Contoh
Nuxt 4 Pages & Config
~50
API Handlers (Drizzle)
~80… See the full description on the dataset page: https://huggingface.co/datasets/syaher/web-builder-dataset.hyundai_web_dataweb-dataset
