datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wisconsin-building-codes-qa
Wisconsin Building Codes Q&A Dataset
Dataset Description
This dataset contains 13200 question-answer pairs focused on Wisconsin building codes, specifically covering:
Building code requirements and regulations
Administrative procedures and enforcement
Construction standards and specifications
Permit processes and compliance
Dataset Structure
Training samples: 11880
Validation samples: 1320
Each sample contains:
instruction: A question about Wisconsin… See the full description on the dataset page: https://huggingface.co/datasets/carlscape/wisconsin-building-codes-qa.building-engineering-synthetic-dataset-v5
Building Engineering Synthetic Dataset (V5)
Repository: Irfanuruchi/building-engineering-synthetic-dataset-v5
This repository contains a synthetic dataset for training engineering reasoning models focused on building engineering calculations and sanity checks.
The dataset was generated using physics-based engineering equations and structured prompts suitable for LLM fine-tuning.
It was used to train:
Irfanuruchi/qwen2.5-1.5b-buildeng-precheck-lora-v5
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/Irfanuruchi/building-engineering-synthetic-dataset-v5.wisconsin-building-codes-grpo
Wisconsin Building Codes Q&A Dataset (GRPO-Formatted)
This dataset is a version of the Wisconsin Building Codes Q&A Dataset formatted specifically for Grouped-Reward-Optimization (GRPO) training with libraries like TRL and unsloth.
Dataset Description
This dataset contains 13,200 prompts designed for training preference models. Each record includes a user prompt (prompt), a "chosen" high-quality response, and a placeholder for a "rejected" response.
Training samples: 11… See the full description on the dataset page: https://huggingface.co/datasets/carlscape/wisconsin-building-codes-grpo.community-building-playbook
Community Building Playbook
Build engaged developer and user communities from scratch — Ambassador programs, CLG, Discord/Slack setup, event ops, health KPIs
English | 中文
Need a 1-on-1 community strategy call? Book a session — Contact @Iris_carrot on Telegram
Or visit gingiris.tools — Iris's full skill library and consulting practice.
What is this?
A battle-tested community building playbook for developer tools and SaaS products. Real… See the full description on the dataset page: https://huggingface.co/datasets/Gingiris/community-building-playbook.Building_External_Networks_Ecosystems_Theory
Building External Networks Ecosystems — Theory
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Building_External_Networks_Ecosystems_Theory.Building_External_Networks_Ecosystems_Practical
Building External Networks Ecosystems — Practical
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Building_External_Networks_Ecosystems_Practical.Building_Trust_Content_1
Building Trust Content 1
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Building_Trust_Content_1.
