gvic-unb/public-health-news-DF-qa
Public Health Brasília QA Corpus Dataset Summary Public Health Brasília QA Corpus is a dataset for evaluating Retrieval-Augmented Generation (RAG) systems over public health news articles from the Secretaria de Saúde do Distrito Federal (SES-DF), Brazil. It consists of two components: a QA evaluation corpus with location-aware question-answer pairs, and a knowledge base corpus of 1,688 public health news articles that serves as the retrieval source for the RAG… See the full description on the dataset page: https://huggingface.co/datasets/gvic-unb/public-health-news-DF-qa.
Public Health Brasília QA Corpus
Dataset Summary
Public Health Brasília QA Corpus is a dataset for evaluating Retrieval-Augmented Generation (RAG) systems over public health news articles from the Secretaria de Saúde do Distrito Federal (SES-DF), Brazil. It consists of two components: a QA evaluation corpus with location-aware question-answer pairs, and a knowledge base corpus of 1,688 public health news articles that serves as the retrieval source for the RAG pipeline.
Questions focus on public health facilities, services, and programs distributed across the administrative regions of the Federal District. QA pairs were generated from SES-DF news articles using Claude Haiku 4.5 (claude-haiku-4-5-20251001), with prompt design guided by a public health expert to simulate the inquiry style of researchers, managers, and citizens.
Questions cover four types:
Dataset Structure
QA Splits (default config)
Each QA instance contains the following fields:
Knowledge Base (knowledge_base config)
Each corpus instance contains the following fields:
Question Type Distribution
The following distribution applies across the full generated pool (before splitting):
Generation and Filtering Pipeline
QA pairs were generated using Claude Haiku 4.5 (claude-haiku-4-5-20251001), prompted in Brazilian Portuguese, and filtered through the following steps:
- Pairs with empty question or answer fields were discarded.
- Each question was required to explicitly contain at least one location name mentioned in the source article (e.g., "Ceilândia"). Pairs that did not satisfy this constraint were removed via case-insensitive substring matching.
In total, 933 QA pairs were initially generated, of which 283 were discarded for failing the locality check (30.3% discard rate), resulting in 650 valid pairs. These were then split into a validation set of 150 samples and a final evaluation set of 500 samples.
Data Examples
QA pair (test split)
{
"id": 1,
"noticia": "UBS da Asa Norte atua na conscientização sobre HIV e câncer de pele",
"localidade_usada": "Asa Norte",
"tipo": "identificação de serviço (qual unidade de saúde, qual programa ou campanha)",
"pergunta": "Qual unidade de saúde da Asa Norte promoveu uma manhã de ações preventivas com testagem para HIV e ISTs em dezembro?",
"resposta": "UBS 1 da Asa Norte"
}Knowledge base article (corpus split)
{
"title": "Dia do Trabalho: confira o funcionamento das unidades de saúde",
"content": "Serviços de urgência e emergência funcionam normalmente...",
"url": "https://www.saude.df.gov.br/...",
"localidades": "taguatinga;asa norte;sao sebastiao;gama",
"tipos_unidade": "hospital"
}Source Corpus
Articles were collected from the public health news portal of the Secretaria de Saúde do Distrito Federal (SES-DF), covering health facilities, vaccination campaigns, epidemiological alerts, and public health programmes across the Federal District.
Languages
Brazilian Portuguese (pt-BR).
Licensing
This dataset is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Citation
If you use this dataset, please cite:
@inproceedings{borges2026integrating,
title = {Integrating Large Language Models and Geospatial Data for Territory-Aware Public Health Surveillance},
author = {Vinicius R. P. Borges and Arthur Soares Magalh{\~a}es Freire and Luis Paulo Faina Garcia and
Anna Carolina Faleiros Martins and Flavio de Barros Vidal and Aleteia Araujo and
Edward Torres Maia and Gabriel Maia Veloso and Wagner de Jesus Martins and
Vladimir Molchanov and Lars Linsen},
booktitle = {Workshop on Foundation Models for Social Good @ International Joint Conference on Artificial Intelligence - European Conference on Artificial Intelligence (IJCAI-ECAI) 2026},
year = {2026},
url = {https://openreview.net/forum?id=qeXMRztxxY}
}Contact
For questions or feedback, please open an issue in the project repository.
