datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aec-bench
AEC-Bench: A Multimodal Dataset for Architecture, Engineering, and Construction
Section
What it covers
Overview
What the dataset contains
Task taxonomy
Scopes, task families, instance counts
Accessing the dataset
manifest.jsonl, prefetching files from URLs
License
Apache 2.0
Citation
BibTeX
Overview
AEC-Bench is a multimodal dataset of real-world Architecture, Engineering, and Construction (AEC) documents — construction drawings, floor… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/aec-bench.AECBench
🏗️ AECBench
🇺🇸 English | 🇨🇳 中文说明
Project Introduction
AECBench is an open-source large language model Architecture, Engineering & Construction (AEC) domain evaluation benchmark jointly released by East China Architectural Design & Research Institute Co., Ltd. (ECADI) of China Construction Group and Tongji University. This dataset aims to systematically evaluate large language models' (LLMs) knowledge mastery, understanding, reasoning… See the full description on the dataset page: https://huggingface.co/datasets/jackluoluo/AECBench.aec-rag-dataset
Lumen-Models: AEC-RAG Dataset
Lumen-Models is the premier conversational dataset designed to fine-tune LLMs and empower RAG (Retrieval-Augmented Generation) systems within the Architecture, Engineering, and Construction (AEC) sector.
This dataset features high-fidelity technical dialogues between a BIM Auditor and a GPT Expert, focused on solving real-world challenges regarding regulatory compliance, complex construction codes, and professional industry standards.
Premium… See the full description on the dataset page: https://huggingface.co/datasets/lumen-models/aec-rag-dataset.
