pamessina/T5FactExtractor
T5FactExtractor — Radiology Fact Extractor
T5FactExtractor is a T5-small sequence-to-sequence model that extracts factual statements from chest X-ray radiology report sentences. Given a sentence, it generates a JSON-like list of short clinical facts that can be embedded, compared, or used in metrics such as CXRFEScore.
It is stage 1 of the Extracting and Encoding framework from Findings of ACL 2024:
- Fact extraction — this model (
pamessina/T5FactExtractor) - Fact encoding — `pamessina/CXRFE`
Model details
Output format
The model generates a string containing a JSON array of fact strings, for example:
["small right pleural effusion", "normal heart size"]Downstream code (including `cxrfescore`) parses that array, deduplicates facts, and lightly cleans repeated words.
How to use
Standalone (Transformers)
import re
import json
import torch
from transformers import T5ForConditionalGeneration, T5TokenizerFast
device = "cuda" if torch.cuda.is_available() else "cpu"
model_id = "pamessina/T5FactExtractor"
tokenizer = T5TokenizerFast.from_pretrained(model_id)
model = T5ForConditionalGeneration.from_pretrained(model_id).to(device)
model.eval()
sentence = "There is a small right pleural effusion. The heart size is normal."
# Prefer one sentence at a time (reports are usually sentence-split first).
inputs = tokenizer(sentence, padding="longest", return_tensors="pt")
input_ids = inputs["input_ids"].to(device)
attention_mask = inputs["attention_mask"].to(device)
with torch.no_grad():
output_ids = model.generate(
input_ids=input_ids,
attention_mask=attention_mask,
max_new_tokens=input_ids.shape[1] * 4,
num_beams=1,
)
raw = tokenizer.batch_decode(output_ids, skip_special_tokens=True)[0]
print("raw:", raw)
# Minimal parse (same idea as cxrfescore.text_utils.parse_facts)
match = re.search(r"\[.*", raw)
if match:
facts_str = match.group()
if not facts_str.endswith("]"):
facts_str += "]"
facts = json.loads(facts_str)
print("facts:", facts)Easiest path: CXRFEScore
For full reports, the package sentence-splits, runs this extractor, aggregates unique facts, and (optionally) embeds them with CXRFE:
pip install cxrfescorefrom cxrfescore import CXRFEScore
metric = CXRFEScore(device="cuda")
reports = [
"There is a small right pleural effusion. The heart size is normal.",
]
facts_per_report = metric.extract_facts(reports)
print(facts_per_report[0])Demo notebook: CXR-Fact-Encoder / notebooks/cxrfescore_demo.ipynb
Related resources
- Paper hub: https://github.com/PabloMessina/CXR-Fact-Encoder
- Metric package: https://github.com/PabloMessina/CXRFEScore · PyPI
- Companion fact encoder: https://huggingface.co/pamessina/CXRFE
- ACL Anthology: https://aclanthology.org/2024.findings-acl.236/
- arXiv: https://arxiv.org/abs/2407.01948
Citation
If you use T5FactExtractor, please cite:
@inproceedings{messina-etal-2024-extracting,
title = "Extracting and Encoding: Leveraging Large Language Models and Medical Knowledge to Enhance Radiological Text Representation",
author = "Messina, Pablo and
Vidal, Rene and
Parra, Denis and
Soto, Alvaro and
Araujo, Vladimir",
booktitle = "Findings of the Association for Computational Linguistics: ACL 2024",
month = aug,
year = "2024",
address = "Bangkok, Thailand",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2024.findings-acl.236/",
doi = "10.18653/v1/2024.findings-acl.236",
pages = "3955--3986"
}