CoolFace
Datasetpublic

gvic-unb/public-health-news-DF-qa

Public Health Brasília QA Corpus Dataset Summary Public Health Brasília QA Corpus is a dataset for evaluating Retrieval-Augmented Generation (RAG) systems over public health news articles from the Secretaria de Saúde do Distrito Federal (SES-DF), Brazil. It consists of two components: a QA evaluation corpus with location-aware question-answer pairs, and a knowledge base corpus of 1,688 public health news articles that serves as the retrieval source for the RAG… See the full description on the dataset page: https://huggingface.co/datasets/gvic-unb/public-health-news-DF-qa.

sourceHugging Facecc-by-4.0updated 27d agoView on Hugging Face
0likes68downloads
Dataset Card

Public Health Brasília QA Corpus

Dataset Summary

Public Health Brasília QA Corpus is a dataset for evaluating Retrieval-Augmented Generation (RAG) systems over public health news articles from the Secretaria de Saúde do Distrito Federal (SES-DF), Brazil. It consists of two components: a QA evaluation corpus with location-aware question-answer pairs, and a knowledge base corpus of 1,688 public health news articles that serves as the retrieval source for the RAG pipeline.

Questions focus on public health facilities, services, and programs distributed across the administrative regions of the Federal District. QA pairs were generated from SES-DF news articles using Claude Haiku 4.5 (claude-haiku-4-5-20251001), with prompt design guided by a public health expert to simulate the inquiry style of researchers, managers, and citizens.

Questions cover four types:

TypeDescription
FactualWho performed it, when it occurred, where it happened
InterpretiveWhy it was necessary, how it works, what the objective is
Service identificationWhich health unit, which programme or campaign
Community impactWhich groups were served, what the territorial coverage was

Dataset Structure

QA Splits (default config)

SplitSizeDescription
test500 pairsMain evaluation set
validation150 pairsUsed for prompt refinement and agent tuning

Each QA instance contains the following fields:

FieldTypeDescription
idintUnique pair identifier
noticiastringTitle of the source news article
localidade_usadastringLocation name explicitly used in the question
tipostringQuestion type (factual, interpretive, service identification, community impact)
perguntastringQuestion in Brazilian Portuguese
respostastringReference answer grounded in the source news article

Knowledge Base (knowledge_base config)

SplitSizeDescription
corpus1,688 articlesRAG knowledge base of SES-DF public health news

Each corpus instance contains the following fields:

FieldTypeDescription
titlestringArticle title
contentstringFull article text
urlstringSource URL
localidadesstringSemicolon-separated list of administrative regions mentioned in the article
tipos_unidadestringType of health facility or service referenced

Question Type Distribution

The following distribution applies across the full generated pool (before splitting):

Question TypeCount
Factual157
Interpretive162
Service identification163
Community impact168
Total650

Generation and Filtering Pipeline

QA pairs were generated using Claude Haiku 4.5 (claude-haiku-4-5-20251001), prompted in Brazilian Portuguese, and filtered through the following steps:

  1. 1.Pairs with empty question or answer fields were discarded.
  2. 2.Each question was required to explicitly contain at least one location name mentioned in the source article (e.g., "Ceilândia"). Pairs that did not satisfy this constraint were removed via case-insensitive substring matching.

In total, 933 QA pairs were initially generated, of which 283 were discarded for failing the locality check (30.3% discard rate), resulting in 650 valid pairs. These were then split into a validation set of 150 samples and a final evaluation set of 500 samples.

Data Examples

QA pair (test split)

json
{
  "id": 1,
  "noticia": "UBS da Asa Norte atua na conscientização sobre HIV e câncer de pele",
  "localidade_usada": "Asa Norte",
  "tipo": "identificação de serviço (qual unidade de saúde, qual programa ou campanha)",
  "pergunta": "Qual unidade de saúde da Asa Norte promoveu uma manhã de ações preventivas com testagem para HIV e ISTs em dezembro?",
  "resposta": "UBS 1 da Asa Norte"
}

Knowledge base article (corpus split)

json
{
  "title": "Dia do Trabalho: confira o funcionamento das unidades de saúde",
  "content": "Serviços de urgência e emergência funcionam normalmente...",
  "url": "https://www.saude.df.gov.br/...",
  "localidades": "taguatinga;asa norte;sao sebastiao;gama",
  "tipos_unidade": "hospital"
}

Source Corpus

Articles were collected from the public health news portal of the Secretaria de Saúde do Distrito Federal (SES-DF), covering health facilities, vaccination campaigns, epidemiological alerts, and public health programmes across the Federal District.

Languages

Brazilian Portuguese (pt-BR).

Licensing

This dataset is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.

Citation

If you use this dataset, please cite:

bibtex
@inproceedings{borges2026integrating,
  title     = {Integrating Large Language Models and Geospatial Data for Territory-Aware Public Health Surveillance},
  author    = {Vinicius R. P. Borges and Arthur Soares Magalh{\~a}es Freire and Luis Paulo Faina Garcia and
               Anna Carolina Faleiros Martins and Flavio de Barros Vidal and Aleteia Araujo and
               Edward Torres Maia and Gabriel Maia Veloso and Wagner de Jesus Martins and
               Vladimir Molchanov and Lars Linsen},
  booktitle = {Workshop on Foundation Models for Social Good @ International Joint Conference on Artificial Intelligence - European Conference on Artificial Intelligence (IJCAI-ECAI) 2026},
  year      = {2026},
  url       = {https://openreview.net/forum?id=qeXMRztxxY}
}

Contact

For questions or feedback, please open an issue in the project repository.