data-driven
cwicr-construction-rates
CWICR — Construction Works, Items, Costs & Resources
A multilingual, machine-readable database of national construction rate books for 30 countries / language locales. Each rate is fully decomposed into its work composition and resource breakdown (labour, machinery, materials), with unit prices, hierarchical classification, and physical parameters preserved in the source language.
This dataset is the tabular source-of-truth behind the cwicr-vector-db-bgem3-v3 Qdrant snapshots. Use… See the full description on the dataset page: https://huggingface.co/datasets/DataDrivenConstruction/cwicr-construction-rates.WolneLektury-TTS-Polish
WolneLektury-TTS-Polish
A large-scale, high-quality Polish speech dataset for text-to-speech and automatic speech recognition.
Data Source
Derived from Wolne Lektury (Free Readings), a Polish digital library with public domain audiobooks featuring professional voice actors.
Dataset Statistics
Metric
Value
Total samples
383,710
Total duration
997 hours
Unique narrators
1207
Male samples
294,756 (767h)
Female samples
88,945 (230h)
Average… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/WolneLektury-TTS-Polish.TTS-German
TTS-German
High-quality German speech dataset for TTS and ASR, derived from CML-TTS German.
Processing Pipeline
Standardize → 24kHz mono WAV, loudness normalize
Transcribe → WhisperX word-level timestamps
Segment → ≤12s at word boundaries
Denoise → DeepFilterNet
Quality filter → DNSMOS ≥ 2.5
G2P → IPA phonemes (custom dictionary)
Statistics
Metric
Value
Samples
670,509
Hours
1250h
Sample rate
24kHz mono
Max duration
12s
Schema… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-German.cwicr-vector-db-bgem3-v3
CWICR Vector Database — BGE-M3 V3 Snapshots
Production Qdrant snapshots for CWICR (Construction Works Items, Costs & Resources) — a multilingual catalogue of construction rate databases covering 30 countries / language locales. Each snapshot encodes one country's rate book using the BAAI/bge-m3 embedder and is ready to restore directly into a Qdrant server for hybrid semantic search.
These snapshots are the V3 production artifacts produced by the OpenConstructionEstimate / CWICR… See the full description on the dataset page: https://huggingface.co/datasets/DataDrivenConstruction/cwicr-vector-db-bgem3-v3.TTS-English-HiFiTTS-English-LibriTTS
