spadeMIA/pmc_finetune_corpus_1024-2040_tokens
PMC 1024-2040 Biomedical Fine-Tuning Corpus Summary This is a cleaned biomedical long-text corpus for autoregressive language-model fine-tuning and held-out evaluation. split rows role train 10,000 fine-tuning test 1,000 held-out evaluation The public schema is text-only: text: string No PMCID, date, license, URL, or provenance fields are included in the public dataset files. Token Contract The corpus is built for… See the full description on the dataset page: https://huggingface.co/datasets/spadeMIA/pmc_finetune_corpus_1024-2040_tokens.
This repository belongs to spadeMIA on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
