datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CFR-Title-2-Uniform-Administrative-Requirements-Cost-Principles-And-Audit
Title 2 CFR Uniform Administrative Requirements, Cost Principles, and Audit Question-Answer Dataset
Dataset Summary
This dataset contains document-grounded question-and-answer samples based on Title 2 of the Code of Federal Regulations—Uniform Administrative Requirements, Cost Principles, and Audit Requirements for Federal Awards, commonly referred to as the Uniform Guidance.
The Uniform Guidance establishes Government-wide requirements for administering Federal… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/CFR-Title-2-Uniform-Administrative-Requirements-Cost-Principles-And-Audit.openai-terra-batch-wiki-brazil-1000-partial-20260724-01
OpenAI Terra Batch — Wikipédia PT-BR (run parcial)
Checkpoint publicável de uma execução real e interrompida do fluxo
document_task_matrix. A execução planejou gerar uma matriz de 1.000
documentos da Wikipédia em português por 25 tasks canônicas usando a Responses
API Batch e o modelo gpt-5.6-terra.
Este repositório não representa a conclusão dos 25.000 pares planejados. Ele
contém somente os 1.282 candidatos aceitos após a reconciliação offline de
todos os resultados Batch já… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/openai-terra-batch-wiki-brazil-1000-partial-20260724-01.roofing-cost-index
US Residential Roofing Cost Index (2026)
Dataset Summary
This dataset contains highly localized, objective residential roof replacement pricing indices for 505 major US cities across all 50 states. All pricing figures represent synthesized, algorithmically compiled estimates for the year 2026 by the Shingle Geek pricing engine.
The dataset provides dual cost models to inject complete transparency into the residential home improvement market:
Fair Contractor… See the full description on the dataset page: https://huggingface.co/datasets/ShingleGeek/roofing-cost-index.wikipedia-pt-br-instruct-5k
Wikipedia PT-BR Instruct
wikipedia-pt-br-instruct is a synthetic supervised fine-tuning (SFT)
dataset in Brazilian Portuguese generated from Wikipedia-derived documents.
This release is an intermediate evaluation dataset produced with the
sft-dataset-creator pipeline from the run
wiki-ptbr-extract-calib-5kdocs-14tasks. It was generated from a fixed
revision of costadev00/wikipedia-pt-br-extract:
cdbd07dc4a3de6e64632c718710b3ae0ebaeb0ff
The dataset is intended for intermediate… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-instruct-5k.smoke-openai-terra-batch-brasil-25-20260724-01
Smoke OpenAI Terra Batch — Brasil × 25 tasks
Run real de validação do fluxo matricial document_task_matrix, executada
sobre um único documento da Wikipédia em português com o título Brasil.
Cada uma das 25 tasks canônicas recebeu exatamente um slot inicial.
Resultado
status: completed
documentos: 1
pares planejados: 25
exemplos aceitos: 25
pares pulados: 0
pares esgotados: 0
resultados reais do backend: 27
retries com nova chamada: 2
backend: openai_api… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/smoke-openai-terra-batch-brasil-25-20260724-01.finops_token_cost_prefix_cache_terminator_teaser
🚀 Enterprise FinOps - AI Token Cost & Prefix-Cache Terminator (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Enterprise FinOps - AI Token Cost & Prefix-Cache Terminator on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
📦 What is Inside the Full Production Package:
500 Verified FAANG v2.0 Scenarios (100% AST-Valid Python)… See the full description on the dataset page: https://huggingface.co/datasets/emgena/finops_token_cost_prefix_cache_terminator_teaser.dolly-15k-rlhf-instructgpt-format
Dolly 15k RLHF Datasets in InstructGPT Format
This repository packages databricks/databricks-dolly-15k into three RLHF-oriented
dataset configurations inspired by the InstructGPT data flow:
sft: supervised fine-tuning examples with prompt, completion, and text.
rm_schema: reward-modeling schema/prompt pool with empty chosen and rejected
fields, reference_response, and ready_for_rm=false.
rm_synthetic: reward-modeling proxy pairs where Dolly reference_response is
used as chosen and… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/dolly-15k-rlhf-instructgpt-format.autonomous-driving-ethical-cost-field-construction-v0.1
What this dataset tests
Whether an intelligence system can constructan ethical cost field for a driving scene.
The task is not to choose an action.The task is to model how harm distributes across agents.
Required outputs
ethical cost field
agent harm vectors
aggregate deformation score
rights infringement index
uncertainty band
Use case
Foundation layer for ethical navigation systems.Trains models to map harm before selecting actions.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-ethical-cost-field-construction-v0.1.pc-insurance-cost-estimator
Property & Casualty insurance dataset
This dataset shows chat with insurance expert where damage to property is mentioned and the assistant
responds with estimate of cost to repair in USD. Dataset has been egnerated using Claude Sonnet 3.5
and GPT4-Omni models.
wikipedia-pt-br-instructions-sft
wikipedia-pt-br-instructions-gemma-alpaca
Dataset de Instruction Following em formato Alpaca puro, com exatamente os campos instruction, input e output.
Origem
Derivado do dataset local instruction_following produzido por wiki-if-builder, por sua vez derivado de costadev00/wikipedia-pt-br-extract.
Campos
instruction: comando em português brasileiro.
input: contexto mínimo opcional.
output: resposta esperada.
Licença e limitações
A licença herdada é… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-instructions-sft.stf-acordaos
STF Acordaos
Dataset de acordaos do Supremo Tribunal Federal em portugues, preparado para uso em NLP juridico e pre-treinamento continuo.
Esta versao deriva de costadev00/stf-acordaos-cpt-2048, mas publica documentos reconstruidos em vez de chunks. O objetivo e oferecer uma fonte sem chunk_index/chunk_total, para que o usuario possa aplicar o proprio chunking conforme o contexto do modelo ou do pipeline.
Origem dos dados
Fonte publica:… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/stf-acordaos.llm-cost-benchmark
LLM Cost Benchmark
Token pricing and latency benchmarks across major LLM providers. Updated dataset for cost estimation, budget planning, and provider comparison.
Fields
provider: API provider (OpenAI, Anthropic, Google, Meta, Mistral)
model: Model name
input_cost_per_1m: Input cost per 1M tokens (USD)
output_cost_per_1m: Output cost per 1M tokens (USD)
context_window: Maximum context window size
max_output: Maximum output tokens
median_latency_ms: Median latency for… See the full description on the dataset page: https://huggingface.co/datasets/zachz/llm-cost-benchmark.ptdbench-reward-design-reward-min-cost-reducing-lnds-020-dataset
PTDBench dataset snapshot: reward_min_cost_reducing_lnds_020
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-min-cost-reducing-lnds-020-dataset.stf-acordaos-cpt-2048
STF Acordaos CPT 2048
Dataset de acórdãos do STF preparado para continuous pre-training (CPT), com foco em textos jurídicos e preservação da cauda final dos documentos longos (parte 2, parte 3, etc.).
Origem dos dados
Fonte pública: https://dadosabertos.c3sl.ufpr.br/acordaos/json/
Arquivos de origem utilizados:
DocumentosAcordaos.json
AcordaosVotos.json
AcordaosRelatorios.json
Construção
Fonte canônica: DocumentosAcordaos.json
Labels incluídos: integra, voto… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/stf-acordaos-cpt-2048.books-gutenberg-project-pt-br
Gutenberg Project TokenWeaver CPT 2048 - Unchunked
This dataset contains full-document rows reconstructed from
costadev00/gutenberg-project-tokenweaver-cpt-2048.
The source dataset mixes reconstructed chunk sequences and singleton chunk rows.
For this unchunked release, rows were grouped by metadata.id; when duplicated
singleton rows were present for the same document, the reconstruction kept the
series with the largest chunk_total. Text was joined with inferred text
overlap.… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/books-gutenberg-project-pt-br.wikibooks-instruct
Wikibooks Instruct pt-BR
Synthetic instruction-tuning dataset in Brazilian Portuguese generated from
cleaned Portuguese Wikibooks text.
The published run contains all accepted examples from
20260630T233835Z-wikibooks-instruct-7fec6755.
Summary
Accepted examples: 106,010
Attempted examples: 248,040
Original target: 127,330 examples
Run status in metadata: partial
Source documents selected: 9,095
Split: train only
Source dataset: costadev00/wikibooks-pt-br-extract… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikibooks-instruct.wiki-brazil
Wiki Brazil
Dataset criado a partir da API da Wikipédia em português.
Conteúdo
O dataset contém o artigo Brasil e os artigos em namespace principal
listados nos links internos desse artigo.
Arquivos
wiki_brazil.jsonl: registros em JSONL.
Schema
title: título resolvido do artigo na Wikipédia.
page_id_: identificador numérico da página na Wikipédia.
text: texto completo extraído em plaintext pela API MediaWiki.
Fonte… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wiki-brazil.wikipedia-pt-br-instructions
wikipedia-pt-br-instructions-gemma
Dataset sintético de Instruction Following em português brasileiro, derivado de artigos da Wikipedia pt-BR.
Origem
Os exemplos derivam de costadev00/wikipedia-pt-br-extract e preservam source_page_id, source_title, source_dataset e a licença herdada cc-by-sa-3.0.
Processo
A geração é inspirada em Alpaca e Self-Instruct: cada artigo válido passa por uma chamada de analista documental que produz candidatos de instrução ancorados… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-instructions.
