datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GenoJEPA-Pretraining
GenoJEPA-Pretraining
This dataset provides the pre-training resources used for GenoJEPA, a genomic representation learning framework based on joint-embedding predictive architecture.
GenoJEPA learns semantic representations of DNA sequences by shifting the optimization target from nucleotide-level reconstruction to latent-space semantic alignment. The pre-training data is used to construct global and local sequence views for self-supervised genomic representation learning.… See the full description on the dataset page: https://huggingface.co/datasets/ChengsenWang/GenoJEPA-Pretraining.chess-roberta-pretraining-sansconfigs:
config_name: default
data_files:
split: train
path: train/*.csv
split: eval
path: eval/*.csv
Detecting-Access-Violations-in-a-LLMs-Pre-Training-Data
Beyond Public Access in LLM Pre-Training Data
The official HuggingFace repository for the paper "Beyond Public Access in LLM Pre-Training Data" by The AI Disclosures Project.
Using a legally obtained dataset of 34 copyrighted O'Reilly Media books, we apply the DE-COP membership inference attack method to investigate whether OpenAI's large language models were trained on copyrighted content without consent.
