semaj83/ctmatch_ir
CTMatch Information Retrieval Dataset This is a dataset of processed clinical trials documents, somehwat of a duplication of that found in datasets/ir_datasets except that these have been preprocessed with ctproc to clean and extract useful fields from the clinical trial documents. Note: They are currently saved as text files because of the downstream task in ctmatch, though in the future they may be converted to .csv. Each .txt file has exactly 374648 lines of corresponding data:… See the full description on the dataset page: https://huggingface.co/datasets/semaj83/ctmatch_ir.
This repository belongs to semaj83 on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
