datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
roofing-cost-index
US Residential Roofing Cost Index (2026)
Dataset Summary
This dataset contains highly localized, objective residential roof replacement pricing indices for 505 major US cities across all 50 states. All pricing figures represent synthesized, algorithmically compiled estimates for the year 2026 by the Shingle Geek pricing engine.
The dataset provides dual cost models to inject complete transparency into the residential home improvement market:
Fair Contractor… See the full description on the dataset page: https://huggingface.co/datasets/ShingleGeek/roofing-cost-index.autonomous-driving-ethical-cost-field-construction-v0.1
What this dataset tests
Whether an intelligence system can constructan ethical cost field for a driving scene.
The task is not to choose an action.The task is to model how harm distributes across agents.
Required outputs
ethical cost field
agent harm vectors
aggregate deformation score
rights infringement index
uncertainty band
Use case
Foundation layer for ethical navigation systems.Trains models to map harm before selecting actions.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-ethical-cost-field-construction-v0.1.llm-cost-benchmark
LLM Cost Benchmark
Token pricing and latency benchmarks across major LLM providers. Updated dataset for cost estimation, budget planning, and provider comparison.
Fields
provider: API provider (OpenAI, Anthropic, Google, Meta, Mistral)
model: Model name
input_cost_per_1m: Input cost per 1M tokens (USD)
output_cost_per_1m: Output cost per 1M tokens (USD)
context_window: Maximum context window size
max_output: Maximum output tokens
median_latency_ms: Median latency for… See the full description on the dataset page: https://huggingface.co/datasets/zachz/llm-cost-benchmark.stf-acordaos-cpt-2048
STF Acordaos CPT 2048
Dataset de acórdãos do STF preparado para continuous pre-training (CPT), com foco em textos jurídicos e preservação da cauda final dos documentos longos (parte 2, parte 3, etc.).
Origem dos dados
Fonte pública: https://dadosabertos.c3sl.ufpr.br/acordaos/json/
Arquivos de origem utilizados:
DocumentosAcordaos.json
AcordaosVotos.json
AcordaosRelatorios.json
Construção
Fonte canônica: DocumentosAcordaos.json
Labels incluídos: integra, voto… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/stf-acordaos-cpt-2048.wikipedia-pt-br-instructions
wikipedia-pt-br-instructions-gemma
Dataset sintético de Instruction Following em português brasileiro, derivado de artigos da Wikipedia pt-BR.
Origem
Os exemplos derivam de costadev00/wikipedia-pt-br-extract e preservam source_page_id, source_title, source_dataset e a licença herdada cc-by-sa-3.0.
Processo
A geração é inspirada em Alpaca e Self-Instruct: cada artigo válido passa por uma chamada de analista documental que produz candidatos de instrução ancorados… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-instructions.
