datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
proteinLM-mixed-pretraining-v1
Pretraining mix for Protein language models
In order to construct a solid pre-training data mixture for protein language models we sample a mix of proteins from 3 sources:
MG_Prot50 (https://huggingface.co/datasets/tattabio/OMG_prot50): meta-genomic proteins created by clustering the Open MetaGenomic dataset (OMG) at 50% sequence identity.
UniRef50: UniProt proteins from across all species clustered to 50% sequences identity, downloaded form UniProt on 31th of March 2025… See the full description on the dataset page: https://huggingface.co/datasets/MichelNivard/proteinLM-mixed-pretraining-v1.amharic-pretraining-corpusAmharic Pretraining Corpus is a large-scale dataset (~103M) for general amharic language pretraining tasks. It consists of diverse text sources, including news articles, books, social media posts, government documents, and web content, all written in Amharic.
You can load the dataset as follows
from datasets import load_dataset
ds = load_dataset("yordanoswuletaw/amharic-pretraining-corpus")
pretraining_datasetcontinued-pretraining-llama-format
Open Paws Continued Pretraining Llama Format
Overview
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Specialized Data
Format: CSV (Comma-separated values)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/continued-pretraining-llama-format.CGM-JEPA-Pretraining
CGM-JEPA Pretraining Corpus
Continuous glucose monitor (CGM) time-series corpus used for self-supervised pretraining of CGM-JEPA, X-CGM-JEPA, GluFormer, and TS2Vec encoders in the paper CGM-JEPA: Learning Consistent Continuous Glucose Monitor Representations via Predictive Self-Supervised Pretraining.
Code: https://github.com/cruiseresearchgroup/CGM-JEPA
Pretraining-only corpus. For the labeled downstream-evaluation cohorts (insulin resistance and β-cell dysfunction classification)… See the full description on the dataset page: https://huggingface.co/datasets/CRUISEResearchGroup/CGM-JEPA-Pretraining.GenoJEPA-Pretraining
GenoJEPA-Pretraining
This dataset provides the pre-training resources used for GenoJEPA, a genomic representation learning framework based on joint-embedding predictive architecture.
GenoJEPA learns semantic representations of DNA sequences by shifting the optimization target from nucleotide-level reconstruction to latent-space semantic alignment. The pre-training data is used to construct global and local sequence views for self-supervised genomic representation learning.… See the full description on the dataset page: https://huggingface.co/datasets/ChengsenWang/GenoJEPA-Pretraining.bugfix_pretrainingChEMBL_v33_pretrainingzulu-pretraining-datasetThis is IsiZulu Pretraining Dataset. The dataset was used to pre-train BafoGPT-3B
Books: Zulu-English Dictionary – A dictionary offering Zulu terms with English definitions, ideal for teaching basic word mappings.
Translation: South African Government Speeches – Official speeches in Zulu, which help the model understand structured Zulu sentences and phrases.
Transcription: Zulu Community Corpus – A collection of transcriptions, exposing the model to real-life conversational Zulu.
Document:… See the full description on the dataset page: https://huggingface.co/datasets/ChallengerSpaceShuttle/zulu-pretraining-dataset.chess-roberta-pretraining-sansconfigs:
config_name: default
data_files:
split: train
path: train/*.csv
split: eval
path: eval/*.csv
Detecting-Access-Violations-in-a-LLMs-Pre-Training-Data
Beyond Public Access in LLM Pre-Training Data
The official HuggingFace repository for the paper "Beyond Public Access in LLM Pre-Training Data" by The AI Disclosures Project.
Using a legally obtained dataset of 34 copyrighted O'Reilly Media books, we apply the DE-COP membership inference attack method to investigate whether OpenAI's large language models were trained on copyrighted content without consent.
pretraining_synthetic_long_100pretraining_synthetic_shortvrepair_pretraining_dataikigai-pretraining-corpuspretraining_synthetic_longpretraining_synthetic_long_100_5_real_examplespretraining_synthetic_long_100_5_real_random_examplespretraining_synthetic_long_100_1_real_random_examplespretraining_synthetic_short_100_placerasbt_pretraininghinglish_pretraining_dataset
