datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
copyright_gpt_neo_1_3Bcopyrightgpt-v1
CopyrightGPT
A fast first-draft video dataset + fingerprint store, built to answer one question:
has this video (or something very similar to it) been seen before?
Each ingested video gets a YouTube-style random id and a folder ("bin") of binary
data files. A perceptual hash (pHash) is computed for sampled frames, so a new
video can be checked against everything already stored to flag likely duplicates /
re-uploads.
This is an early, intentionally simple draft — perceptual-hash… See the full description on the dataset page: https://huggingface.co/datasets/hdcli/copyrightgpt-v1.harmbench_copyright_classifier_hashes
HarmBench Copyright Classifier Hashes
Original files: https://github.com/centerforaisafety/HarmBench/tree/main/data/copyright_classifier_hashes
harmbench_copyright_classifier_hashes
HarmBench Copyright Classifier Hashes
Original files: https://github.com/centerforaisafety/HarmBench/tree/main/data/copyright_classifier_hashes
copyright_mainablation-copyright
catshift
System Requirements
pip install -r requirements.txt
Other Requirements
To run CatShift algorithm, you will need the download models (e.g Pythia-410m).
As well as the GPU of:
NVIDIA A800
VRAM: 80GB
Software Requirement
Ensure the following software is installed before you proceed with the installation of required Python dependencies and execution of the source code:
Python: It is recommended to use version 3.9 or higher.
pip or conda:… See the full description on the dataset page: https://huggingface.co/datasets/bxiong/ablation-copyright.copyright_helmcopyright_unlearningcopyright-traps
Copyright Traps
Copyright traps (see Meeus et al. (ICML 2024)) are unique, synthetically generated sequences
who have been included into the training dataset of CroissantLLM.
This dataset allows for the evaluation of Membership Inference Attacks (MIAs) using CroissantLLM as target model,
where the goal is to infer whether a certain trap sequence was either included in or excluded from the training data.
This dataset contains non-member (label=0) and member (label=1) trap… See the full description on the dataset page: https://huggingface.co/datasets/imperial-cpg/copyright-traps.kl3m-filter-data-dotgov-www.copyright.govkl3m-data-dotgov-www.copyright.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.copyright.gov.harmbench-copyright-hashes
HarmBench copyright — prompts and reference hashes
Everything here is generated from the public
centerforaisafety/HarmBench release by
scripts/build_copyright_hash_parquet.py in eval-framework-companion. Nothing is copied from any
other host.
Contents
File
What
copyright/train-00000-of-00001.parquet
The 100 copyright behaviors (prompt, tags). train is the Hugging Face split name; these are evaluation samples, nothing is trained on them… See the full description on the dataset page: https://huggingface.co/datasets/ividal/harmbench-copyright-hashes.copyright-traps-extra-non-memberscopyrightBooks
ForgetRetainBooks
This dataset is derived from the NarrativeQA dataset, created by Kocisky et al. (2018). NarrativeQA is a dataset for evaluating reading comprehension and narrative understanding.
This dataset is an extraction of the book content from the original NarrativeQA dataset.
Citation
If you want to use this dataset, please also cite the original NarrativeQA dataset.
@article{narrativeqa,
author = {Tom\'a\v s Ko\v cisk\'y and Jonathan Schwarz and Phil Blunsom and… See the full description on the dataset page: https://huggingface.co/datasets/kqwang/copyrightBooks.no-copyright-duwalicopyrightQA
CopyrightQA
This dataset is derived from the NarrativeQA dataset, created by Kocisky et al. (2018). NarrativeQA is a dataset for evaluating reading comprehension and narrative understanding.
This dataset is an extraction of the question answer pairs from the original NarrativeQA dataset. It's original use is to evaluate LLMs forgetting ability using TOFU, created by Maini et al. (2024). TOFU is a benchmark for evaluating unlearning performance of LLMs on realistic tasks.… See the full description on the dataset page: https://huggingface.co/datasets/kqwang/copyrightQA.copyrighted-books-are-publicdeepseek-llm-7b-chat-2-29-kyssen-164-kyssen-copyright-outputsdeepseek-llm-7b-chat-kyssen-copyright-outputsbhl_copyright_statuses_classifiedcolabel-copyright-substitution-riskfinewebs-copyright-domains
List of domains that were removed from FineWeb(-2)
An easy-access dataset that contains the domains that were removed from FineWeb(-2) after take-down requests. This should allow others to quickly filter out domains that should not be included in data collection, removing redundant data processing. If used in this way, as a domain-wide filter, it should make life easier for the copyright holders, too, who will not have to resubmit cease-and-desists for each new webcrawl.
Currently… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/finewebs-copyright-domains.Meta-Llama-3-8B-Instruct-copyright-kyssen-stage1-29-2-164-violations-12-copyrightTorrent_Copyright_Mappingsbhl_copyright_statuses
Biodiversity Heritage Library Copyright Statuses
This dataset contains all unique copyright statuses present in the items.txt.gz file of the Biodiversity Heritage Library open dataset on AWS Open Data. The unique copyright statuses were extracted, grouped and sorted by frequency using the following DuckDB query:
COPY (SELECT CopyrightStatus, COUNT(*) as Count FROM read_csv('https://bhl-open-data.s3.amazonaws.com/data/item.txt.gz') GROUP BY CopyrightStatus ORDER BY Count DESC) TO… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/bhl_copyright_statuses.gemma-2-9b-it-copyright-33-copyright-outputsMeta-Llama-3-8B-Instruct-copyright-33-copyright-with-books-100-kyssen-copyright-outputscopyright-safetygemma-2-9b-it-copyright-outputsMeta-Llama-3-8B-Instruct-copyright-33-copyright-with-books-100-copyright-outputs
