datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rephrased-web-data-quality-study
Rephrased Web Data Quality Study
LLM-as-judge evaluation of ~4,000 examples from HuggingFaceFW/finephrase (1,000 sampled per split, 86 dropped due to judge parse failures, 3,914 successfully evaluated).
Judge: Claude Sonnet 4.6 via OpenRouter | Cost: ~$45
Quality Scores (1-5 scale)
Metric
FAQ (n=965)
Table (n=979)
Tutorial (n=976)
Math (n=994)
Faithfulness
1.82
1.72
1.90
1.49
Info preservation
1.93
1.64
1.99
1.47
Appropriateness
3.54
2.87
2.48
1.67… See the full description on the dataset page: https://huggingface.co/datasets/ratishsp/rephrased-web-data-quality-study.lexi-coding-web-clean-datasets
lexi-coding-web-clean-datasets
lexi-coding-web-clean-datasets by Reallexi LLC AI Model Builder — llm.reallexi.io
Copyright (c) 2026 Reallexi LLC. All rights reserved.
A retrieval index built by Reallexi AI Model Builder: source text chunked, embedded, and stored for nearest-neighbor retrieval. This is not a causal-language-model checkpoint and cannot be loaded with AutoModelForCausalLM.
Contents
Source data… See the full description on the dataset page: https://huggingface.co/datasets/reallexi/lexi-coding-web-clean-datasets.bop-webdataset-shards
