datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
so-render-dpo
SO-RENDER: Stack Overflow Rendering Preference Dataset
A preference dataset built from Stack Overflow questions about web rendering — specifically the shift from server-side rendering (SSR) to client-side SPAs and back to SSR again. The dataset is designed for training and evaluating language models on data where community preferences change over time.
The story behind the data
Web development went through a pretty clear cycle over the past decade or so:
Early era… See the full description on the dataset page: https://huggingface.co/datasets/basab1142/so-render-dpo.so-render-dpo-v3
SO-RENDER: Stack Overflow Rendering Preference Dataset
A preference dataset built from Stack Overflow questions about web rendering — specifically the shift from server-side rendering (SSR) to client-side SPAs and back to SSR again. The dataset is designed for training and evaluating language models on data where community preferences change over time.
The story behind the data
Web development went through a pretty clear cycle over the past decade or so:
Early era… See the full description on the dataset page: https://huggingface.co/datasets/basab1142/so-render-dpo-v3.meaningfulness-cross-language-rendering
Cross-Language Rendering for Meaning vs Meaningfulness (Paper B 2026ap)
HF dataset DOI: 10.57967/hf/8971
Companion paper concept DOI: 10.5281/zenodo.20409701
Companion GitHub mirror: https://github.com/spectralbranding/meaningfulness-papers/tree/main/meaning-meaningfulness-empirical
Dataset Summary
This dataset contains the multi-language rendering and extraction artifacts demonstrating Proposition P4 (rendering-equivalence under spine-preservation) from Zharnikov… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/meaningfulness-cross-language-rendering.so-render-dpo-v4
SO-RENDER: Stack Overflow Rendering Preference Dataset
A preference dataset built from Stack Overflow questions about web rendering — specifically the shift from server-side rendering (SSR) to client-side SPAs and back to SSR again. The dataset is designed for training and evaluating language models on data where community preferences change over time.
The story behind the data
Web development went through a pretty clear cycle over the past decade or so:
Early era… See the full description on the dataset page: https://huggingface.co/datasets/basab1142/so-render-dpo-v4.svg-chart-render-v1
SVG Chart Render Mix v1
Training data for fine-tuning a small code model (DeepSeek Coder 1.3B)
to map (chart specification JSON) to inline SVG code.
Part of the SQL Agent LLMOps project.
Total
Sources
Input
Output
~25,000 rows
2
structured JSON chart spec
rendered SVG string
Part of the SQL Agent LLMOps project
Dataset
Model
Role
DanielRegaladoCardoso/text-to-sql-mix-v2
Qwen 2.5 Coder 7B
NL question to SQL… See the full description on the dataset page: https://huggingface.co/datasets/DanielRegaladoCardoso/svg-chart-render-v1.PDD3_text_rendered_v2
