CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MichelNivard /proteinLM-mixed-pretraining-v1 Pretraining mix for Protein language models In order to construct a solid pre-training data mixture for protein language models we sample a mix of proteins from 3 sources: MG_Prot50 (https://huggingface.co/datasets/tattabio/OMG_prot50): meta-genomic proteins created by clustering the Open MetaGenomic dataset (OMG) at 50% sequence identity. UniRef50: UniProt proteins from across all species clustered to 50% sequences identity, downloaded form UniProt on 31th of March 2025… See the full description on the dataset page: https://huggingface.co/datasets/MichelNivard/proteinLM-mixed-pretraining-v1.texttoken-classification100M<n<1B0 likes354 downloads1y agoHugging Face02yordanoswuletaw /amharic-pretraining-corpusAmharic Pretraining Corpus is a large-scale dataset (~103M) for general amharic language pretraining tasks. It consists of diverse text sources, including news articles, books, social media posts, government documents, and web content, all written in Amharic. You can load the dataset as follows from datasets import load_dataset ds = load_dataset("yordanoswuletaw/amharic-pretraining-corpus") texttext-generation100M<n<1B4 likes293 downloads2y agoHugging Face03Gunulhona /pretraining_datasettext100K<n<1M0 likes266 downloads3y agoHugging Face04open-paws /continued-pretraining-llama-format Open Paws Continued Pretraining Llama Format Overview This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Specialized Data Format: CSV (Comma-separated values) Languages: Multilingual (primarily English) Focus: Animal advocacy and ethical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/continued-pretraining-llama-format.texttext-generation10K<n<100K2 likes63 downloads1y agoHugging Face05CRUISEResearchGroup /CGM-JEPA-Pretraining CGM-JEPA Pretraining Corpus Continuous glucose monitor (CGM) time-series corpus used for self-supervised pretraining of CGM-JEPA, X-CGM-JEPA, GluFormer, and TS2Vec encoders in the paper CGM-JEPA: Learning Consistent Continuous Glucose Monitor Representations via Predictive Self-Supervised Pretraining. Code: https://github.com/cruiseresearchgroup/CGM-JEPA Pretraining-only corpus. For the labeled downstream-evaluation cohorts (insulin resistance and β-cell dysfunction classification)… See the full description on the dataset page: https://huggingface.co/datasets/CRUISEResearchGroup/CGM-JEPA-Pretraining.texttime-series-forecasting100K<n<1M0 likes63 downloads4mo agoHugging Face06ChengsenWang /GenoJEPA-Pretraining GenoJEPA-Pretraining This dataset provides the pre-training resources used for GenoJEPA, a genomic representation learning framework based on joint-embedding predictive architecture. GenoJEPA learns semantic representations of DNA sequences by shifting the optimization target from nucleotide-level reconstruction to latent-space semantic alignment. The pre-training data is used to construct global and local sequence views for self-supervised genomic representation learning.… See the full description on the dataset page: https://huggingface.co/datasets/ChengsenWang/GenoJEPA-Pretraining.tabularn<1K0 likes47 downloads24d agoHugging Face07xin1997 /bugfix_pretrainingtext100K<n<1M2 likes37 downloads3y agoHugging Face08jonghyunlee /ChEMBL_v33_pretrainingtext1M<n<10M0 likes37 downloads3y agoHugging Face09ChallengerSpaceShuttle /zulu-pretraining-datasetThis is IsiZulu Pretraining Dataset. The dataset was used to pre-train BafoGPT-3B Books: Zulu-English Dictionary – A dictionary offering Zulu terms with English definitions, ideal for teaching basic word mappings. Translation: South African Government Speeches – Official speeches in Zulu, which help the model understand structured Zulu sentences and phrases. Transcription: Zulu Community Corpus – A collection of transcriptions, exposing the model to real-life conversational Zulu. Document:… See the full description on the dataset page: https://huggingface.co/datasets/ChallengerSpaceShuttle/zulu-pretraining-dataset.texttext-generationn<1K3 likes19 downloads2y agoHugging Face10TannerGladson /chess-roberta-pretraining-sansconfigs: config_name: default data_files: split: train path: train/*.csv split: eval path: eval/*.csv tabular100M<n<1B0 likes14 downloads2y agoHugging Face11Disclosures-SSRC /Detecting-Access-Violations-in-a-LLMs-Pre-Training-Data Beyond Public Access in LLM Pre-Training Data The official HuggingFace repository for the paper "Beyond Public Access in LLM Pre-Training Data" by The AI Disclosures Project. Using a legally obtained dataset of 34 copyrighted O'Reilly Media books, we apply the DE-COP membership inference attack method to investigate whether OpenAI's large language models were trained on copyrighted content without consent. tabular100K<n<1M0 likes13 downloads10mo agoHugging Face12Gaoj124 /pretraining_synthetic_long_100textn<1K0 likes11 downloads3y agoHugging Face13Gaoj124 /pretraining_synthetic_shorttextn<1K0 likes10 downloads3y agoHugging Face14xin1997 /vrepair_pretraining_datatext100K<n<1M0 likes8 downloads3y agoHugging Face15ThomasTheMaker /ikigai-pretraining-corpustext10K<n<100K0 likes8 downloads10mo agoHugging Face16Gaoj124 /pretraining_synthetic_longtextn<1K0 likes7 downloads3y agoHugging Face17Gaoj124 /pretraining_synthetic_long_100_5_real_examplestextn<1K0 likes3 downloads3y agoHugging Face18Gaoj124 /pretraining_synthetic_long_100_5_real_random_examplestextn<1K0 likes3 downloads3y agoHugging Face19Gaoj124 /pretraining_synthetic_long_100_1_real_random_examplestextn<1K0 likes3 downloads3y agoHugging Face20Gaoj124 /pretraining_synthetic_short_100_placetextn<1K0 likes3 downloads3y agoHugging Face21Haxirus /rasbt_pretrainingtext10K<n<100K0 likes3 downloads2y agoHugging Face22ananttrivedi /hinglish_pretraining_datasettext100K<n<1M0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.