CoolFace
Datasetpublic

HealthDataHub/PARHAF-biomarkers-annotated

Dataset Card for PARHAF-biomarkers-annotated Reporting Issues & Contributing If you encounter any errors or inconsistencies in this dataset, please report them in the discussion section of the "Community" tab on Hugging Face. For more substantial contributions or collaboration opportunities, feel free to contact us directly. Dataset Summary PARHAF-biomarkers-annotated is a subpart of the PARHAF corpus, an open French corpus of… See the full description on the dataset page: https://huggingface.co/datasets/HealthDataHub/PARHAF-biomarkers-annotated.

sourceHugging Facecc-by-4.0updated 6mo agoView on Hugging Face
3likes100downloads
Dataset Card

Dataset Card for PARHAF-biomarkers-annotated

<div align="center"> <p> <a href="https://huggingface.co/spaces/HealthDataHub/PARTAGES" style="display:inline;"><img src="img/PARTAGES BASELINE_RVB.png" alt="Logo" style="height:20px;vertical-align:middle; margin-right:8px; display:inline; margin-bottom:0.2em; margin-top:0.2em;" title="PARTAGES project"; /></a> <a href="https://www.etalab.gouv.fr/wp-content/uploads/2018/11/open-licence.pdf" style="display:inline;" title="Etalab 2.0 license"><img src="img/Logo-licence-ouverte2.svg" style="height:20px; display:inline; margin-bottom:0.2em; margin-top:0.2em;"/></a> <a href="https://creativecommons.org/licenses/by/4.0/deed.en" style="display:inline;"><img src="https://mirrors.creativecommons.org/presskit/buttons/88x31/png/by.png" style="height:20px; display:inline; margin-bottom:0.2em; margin-top:0.2em;" title="CC BY 4.0 license"/></a> </p> </div>

Reporting Issues & Contributing

If you encounter any errors or inconsistencies in this dataset, please report them in the discussion section of the "Community" tab on Hugging Face.

For more substantial contributions or collaboration opportunities, feel free to contact us directly.

Dataset Description

  • —Points of Contact: FILORI Quentin, KHALIL Youness

Dataset Summary

PARHAF-biomarkers-annotated is a subpart of the PARHAF corpus, an open French corpus of human-authored clinical reports of fictional patients.

It was created to support the development and evaluation of clinical NLP systems for the extraction and characterization of genomic and tissue biomarkers from unstructured pathology reports in oncology..

This dataset contains training data only. The test set will remain under embargo to enable future evaluations under controlled conditions, limiting the risk of LLM contamination through prior data exposure. Please contact us for access to the test data.

This training dataset is divided into a train split (80%) and a dev split (20%) to facilitate experimental design and reproducibility across teams. Teams are free to use the full training set or define a different split configuration.

Each patient record was:

  • —written by a senior medical resident
  • —reviewed by another senior medical resident, from the same specialty
  • —annotated by a specialist of the use case
  • —curated by another specialist of the use case
Data statistics

DATASET SUMMARY

IndicatorValue
Complete Dataset
Number of JSON files152
Total annotations2609
Average document length2348 characters
80% Threshold (annotations)2087

Complete dataset distribution by type

TypeAnnotations% of total
LayersRelation91134.92%
SpanBiomarker152858.57%
SpanResultZone1706.52%

Complete dataset distribution by populated field

Type.FieldOccurrences% of total annot.
LayersRelation.Relation91134.92%
SpanBiomarker.AmplificationState56421.62%
SpanBiomarker.Biomarker99438.10%
SpanBiomarker.BiomarkerAttribute75729.01%
SpanBiomarker.MutationBinary311.19%
SpanBiomarker.MutationType70.27%
SpanBiomarker.Normalized99438.10%
SpanBiomarker.Presence99438.10%
SpanResultZone.ResultZone1706.52%

Train / Dev Split Results

IndicatorTRAINDEV
Number of files12131
Number of annotations2101508
Percentage of dataset80.53%19.47%
Avg. document length (chars)22682657

Train / Dev distribution by type

TypeTRAINDEV% train of type
LayersRelation73317880.46%
SpanBiomarker123329580.69%
SpanResultZone1353579.41%

Train / Dev distribution by populated field

Type.FieldTRAINDEV% train of field
LayersRelation.Relation73317880.46%
SpanBiomarker.AmplificationState45111379.96%
SpanBiomarker.Biomarker80518980.99%
SpanBiomarker.BiomarkerAttribute61114680.71%
SpanBiomarker.MutationBinary27487.10%
SpanBiomarker.MutationType5271.43%
SpanBiomarker.Normalized80518980.99%
SpanBiomarker.Presence80518980.99%
SpanResultZone.ResultZone1353579.41%

Train / Dev distribution by value

Type.Field.ValueTRAINDEV% train
SpanBiomarker.Biomarker.TissueBiomarker76418580.51%
LayersRelation.Relation.RefersTo73317880.46%
SpanBiomarker.Presence.True53212181.47%
SpanBiomarker.BiomarkerAttribute.AmplificationState45111379.96%
SpanBiomarker.AmplificationState.Pos2856282.13%
SpanBiomarker.Presence.False2736880.06%
SpanBiomarker.AmplificationState.Neg1665176.50%
SpanResultZone.ResultZone.ResultZone1353579.41%
SpanBiomarker.BiomarkerAttribute.AmplificationRate1152383.33%
SpanBiomarker.Normalized.NoNormalizedNameFound58986.57%
SpanBiomarker.Normalized.KI67411080.39%
SpanBiomarker.Normalized.PDL1371078.72%
SpanBiomarker.Biomarker.GenomicBiomarker41491.11%
SpanBiomarker.Normalized.TTF136881.82%
SpanBiomarker.Normalized.CK738588.37%
SpanBiomarker.Normalized.RE29487.88%
SpanBiomarker.BiomarkerAttribute.MutationBinary27487.10%
SpanBiomarker.Normalized.CK2027487.10%
SpanBiomarker.Normalized.RP25486.21%
SpanBiomarker.Normalized.P5325292.59%
SpanBiomarker.Normalized.P4020676.92%
SpanBiomarker.Normalized.PMS219676.00%
SpanBiomarker.Normalized.CD2019676.00%
SpanBiomarker.Normalized.HER2IHC23195.83%
SpanBiomarker.Normalized.MSH217673.91%
SpanBiomarker.Normalized.MSH617673.91%
SpanBiomarker.Normalized.MLH116672.73%
SpanBiomarker.MutationBinary.True18481.82%
SpanBiomarker.Normalized.TPS15768.18%
SpanBiomarker.Normalized.MSI17480.95%
SpanBiomarker.Normalized.ALK16576.19%
SpanBiomarker.Normalized.CD1012763.16%
SpanBiomarker.Normalized.ROS114573.68%
SpanBiomarker.Normalized.GATA310758.82%
SpanBiomarker.BiomarkerAttribute.MutationName13476.47%
SpanBiomarker.Normalized.CD59756.25%
SpanBiomarker.Normalized.BCL29756.25%
SpanBiomarker.Normalized.P6314287.50%
SpanBiomarker.Normalized.PAX811378.57%
SpanBiomarker.Normalized.CD310376.92%
SpanBiomarker.Normalized.P1612192.31%
SpanBiomarker.Normalized.CYCLINE9375.00%
SpanBiomarker.Normalized.CKAE1ouAE310283.33%
SpanBiomarker.Normalized.BCL66554.55%
SpanBiomarker.Normalized.CK5ou65550.00%
SpanBiomarker.Normalized.MUM16460.00%
SpanBiomarker.Normalized.CDX2100100.00%
SpanBiomarker.Normalized.CHROMOGRANINE8280.00%
SpanBiomarker.MutationBinary.False90100.00%
SpanBiomarker.Normalized.SYNAPTOPHYSINE8188.89%
SpanBiomarker.Normalized.EMA80100.00%
SpanBiomarker.Normalized.CPS7187.50%
SpanBiomarker.Normalized.CD79A70100.00%
SpanBiomarker.BiomarkerAttribute.MutationType5271.43%
SpanBiomarker.Normalized.WT14266.67%
SpanBiomarker.Normalized.CD233350.00%
SpanBiomarker.Normalized.BRAF60100.00%
SpanBiomarker.Normalized.NKX3ou160100.00%
SpanBiomarker.Normalized.EGFR4266.67%
SpanBiomarker.Normalized.CD565183.33%
SpanBiomarker.Normalized.CADHERINE50100.00%
SpanBiomarker.Normalized.PAX550100.00%
SpanBiomarker.Normalized.CD3450100.00%
SpanBiomarker.Normalized.KL150100.00%
SpanBiomarker.Normalized.BEREP440100.00%
SpanBiomarker.Normalized.PS10040100.00%
SpanBiomarker.MutationType.Substitution40100.00%
SpanBiomarker.Normalized.NAPSINE A3175.00%
SpanBiomarker.Normalized.TDT40100.00%
SpanBiomarker.Normalized.CALRETININE30100.00%
SpanBiomarker.Normalized.CK1930100.00%
SpanBiomarker.Normalized.CD1530100.00%
SpanBiomarker.Normalized.SMAD42166.67%
SpanBiomarker.Normalized.ActineMuscleLisse2166.67%
SpanBiomarker.Normalized.CD4530100.00%
SpanBiomarker.Normalized.CD7930100.00%
SpanBiomarker.Normalized.MYELOPEROXYDASE30100.00%
SpanBiomarker.Normalized.VIMENTINE30100.00%
SpanBiomarker.Normalized.D24020100.00%
SpanBiomarker.Normalized.AntiHepatocytes1150.00%
SpanBiomarker.Normalized.PSA20100.00%
SpanBiomarker.Normalized.ACE20100.00%
SpanBiomarker.Normalized.CD3020100.00%
SpanBiomarker.Normalized.HP20100.00%
SpanBiomarker.Normalized.DBA4420100.00%
SpanBiomarker.Normalized.HNF1BETA20100.00%
SpanBiomarker.Normalized.MET20100.00%
SpanBiomarker.Normalized.HER2FISH1150.00%
SpanBiomarker.Normalized.SATB220100.00%
SpanBiomarker.Normalized.CD13820100.00%
SpanBiomarker.Normalized.CALDESMONE20100.00%
SpanBiomarker.MutationType.Deletion020.00%
SpanBiomarker.Normalized.CD21020.00%
SpanBiomarker.Normalized.BAP110100.00%
SpanBiomarker.Normalized.EBNA210100.00%
SpanBiomarker.Normalized.HMB4510100.00%
SpanBiomarker.Normalized.PRAME10100.00%
SpanBiomarker.Normalized.POLE10100.00%
SpanBiomarker.Normalized.GLYPICAN 310100.00%
SpanBiomarker.Normalized.CD3110100.00%
SpanBiomarker.Normalized.BRCA110100.00%
SpanBiomarker.Normalized.BRCA210100.00%
SpanBiomarker.MutationType.OtherType10100.00%
SpanBiomarker.Normalized.RA10100.00%
SpanBiomarker.Normalized.INI110100.00%
SpanBiomarker.Normalized.GFAP10100.00%
SpanBiomarker.Normalized.CD11710100.00%
Data Origin

The clinical reports are extracted from the PARHAF corpus. Please refer to PARHAF documentation for more information about this corpus.

Languages

  • —fr_FR

Dataset Structure

We distribute both a Hugging Face dataset and a standalone version of the corpus. The standalone dataset consists of a JSON file per patient report, in UIMA CAS JSON format. This format constitutes the canonical version of the corpus. The Hugging Face dataset (Parquet/Arrow) is a derived representation generated automatically from the JSON files.

Both formats therefore contain identical information and differ only in storage layout.

One dataset instance corresponds to one report.

Hugging Face dataset

This snippet shows how to extract and iterate over medical report information per patient using the datasets library.

python
import pandas as pd
from datasets import load_dataset

dfs = {cfg: load_dataset("HealthDataHub/PARHAF-biomarkers-annotated", cfg, split="train").to_pandas()
       for cfg in ["document_metadata", "spans", "relations"]}

for patient_raw in dfs["document_metadata"].itertuples():
    report_id = patient_raw.report
    text = patient_raw.full_text
    report_spans = dfs["spans"][dfs["spans"]["report"] == report_id]
    report_relations = dfs["relations"][dfs["relations"]["report"] == report_id]
    ...  

Data Fields

PathTypeDescriptionPossible values
document_metadata
reportstringIdentifiant unique du rapport
full_textstringTexte intégral du rapport
spans
reportstringIdentifiant du rapport
span_idintegerIdentifiant de l'annotation
span_typestringType de l'entité annotéeSpanBiomarker, SpanResultZone
beginintegerOffset de début
endintegerOffset de fin
span_textstringTexte de l'entité
attribute_BiomarkerstringBiomarkerGenomicBiomarker, TissueBiomarker
attribute_PresencestringPresenceFalse, True
attribute_NormalizedstringNormalizedACE, ALK, ActineMuscleLisse, AntiHepatocytes, BAP1, BCL2, BCL6, BEREP4, BRAF, BRCA1
attribute_BiomarkerAttributestringBiomarkerAttributeAmplificationRate, AmplificationState, MutationBinary, MutationName, MutationType
attribute_AmplificationStatestringAmplificationStateNeg, Pos
attribute_ResultZonestringResultZoneResultZone
attribute_MutationBinarystringMutationBinaryFalse, True
attribute_MutationTypestringMutationTypeOtherType, Substitution
relations
reportstringIdentifiant du rapport
relation_idintegerIdentifiant de la relation
source_term_idintegerID du terme source (Dependent)
source_textstringTexte du terme source
target_term_idintegerID du terme cible (Governor)
target_textstringTexte du terme cible
attribute_RelationstringRelationRefersTo

Data Splits

Only the training set is released here. The remaining portion of the corpus will be temporarily embargoed to enable future evaluations under controlled conditions, thereby limiting the risk of large language model contamination through prior exposure to the data. You can evaluate your system on the test set through the CodaBench platform.

Annotation Guidelines

You can find the detailed annotation protocol here: annotation_guidelines.pdf

Licensing Information

This dataset is released under licenses:

  • —CC BY 4.0
  • —Etalab 2.0

Citation Information

[More Information Needed]