CoolFace
Datasetpublic

omniomni/omni-science

GitHub   Website   Paper (Coming Soon) Dataset Details This dataset is a combination of corpora of text from scientific Wikipedia articles and scientific papers across major fields of science. This dataset contains continued-pretrain data. Sources This dataset was sourced from the following open-sourced datasets: Science zeroshot/arxiv-biology legacy-datasets/wikipedia bisectgroup/PubMed_TA… See the full description on the dataset page: https://huggingface.co/datasets/omniomni/omni-science.

sourceHugging Facecc-by-sa-4.0updated 1y agoView on Hugging Face
1likes50downloads

omniomni/omni-science · main · files are served by the source, never re-hosted here