CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bxiong /copyright_gpt_neo_1_3B0 likes5.4k downloads1y agoHugging Face02hdcli /copyrightgpt-v1 CopyrightGPT A fast first-draft video dataset + fingerprint store, built to answer one question: has this video (or something very similar to it) been seen before? Each ingested video gets a YouTube-style random id and a folder ("bin") of binary data files. A perceptual hash (pHash) is computed for sampled frames, so a new video can be checked against everything already stored to flag likely duplicates / re-uploads. This is an early, intentionally simple draft — perceptual-hash… See the full description on the dataset page: https://huggingface.co/datasets/hdcli/copyrightgpt-v1.videovideo-classification0 likes2.5k downloads27d agoHugging Face03mariagrandury /harmbench_copyright_classifier_hashes HarmBench Copyright Classifier Hashes Original files: https://github.com/centerforaisafety/HarmBench/tree/main/data/copyright_classifier_hashes 0 likes325 downloads1y agoHugging Face04swiss-ai /harmbench_copyright_classifier_hashes HarmBench Copyright Classifier Hashes Original files: https://github.com/centerforaisafety/HarmBench/tree/main/data/copyright_classifier_hashes 1 likes257 downloads1y agoHugging Face05bxiong /copyright_main0 likes171 downloads2y agoHugging Face06bxiong /ablation-copyright catshift System Requirements pip install -r requirements.txt Other Requirements To run CatShift algorithm, you will need the download models (e.g Pythia-410m). As well as the GPU of: NVIDIA A800 VRAM: 80GB Software Requirement Ensure the following software is installed before you proceed with the installation of required Python dependencies and execution of the source code: Python: It is recommended to use version 3.9 or higher. pip or conda:… See the full description on the dataset page: https://huggingface.co/datasets/bxiong/ablation-copyright.0 likes136 downloads4mo agoHugging Face07lighteval /copyright_helmtext10K<n<100K0 likes126 downloads1y agoHugging Face08boyiwei /copyright_unlearningtextquestion-answering1K<n<10K0 likes123 downloads2y agoHugging Face09imperial-cpg /copyright-traps Copyright Traps Copyright traps (see Meeus et al. (ICML 2024)) are unique, synthetically generated sequences who have been included into the training dataset of CroissantLLM. This dataset allows for the evaluation of Membership Inference Attacks (MIAs) using CroissantLLM as target model, where the goal is to infer whether a certain trap sequence was either included in or excluded from the training data. This dataset contains non-member (label=0) and member (label=1) trap… See the full description on the dataset page: https://huggingface.co/datasets/imperial-cpg/copyright-traps.tabular1K<n<10K0 likes59 downloads2y agoHugging Face10alea-institute /kl3m-filter-data-dotgov-www.copyright.govtext10K<n<100K0 likes51 downloads2y agoHugging Face11alea-institute /kl3m-data-dotgov-www.copyright.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.copyright.gov.text10K<n<100K1 likes35 downloads1y agoHugging Face12ividal /harmbench-copyright-hashes HarmBench copyright — prompts and reference hashes Everything here is generated from the public centerforaisafety/HarmBench release by scripts/build_copyright_hash_parquet.py in eval-framework-companion. Nothing is copied from any other host. Contents File What copyright/train-00000-of-00001.parquet The 100 copyright behaviors (prompt, tags). train is the Hugging Face split name; these are evaluation samples, nothing is trained on them… See the full description on the dataset page: https://huggingface.co/datasets/ividal/harmbench-copyright-hashes.textn<1K0 likes28 downloads2mo agoHugging Face13imperial-cpg /copyright-traps-extra-non-memberstabular10K<n<100K0 likes23 downloads2y agoHugging Face14kqwang /copyrightBooks ForgetRetainBooks This dataset is derived from the NarrativeQA dataset, created by Kocisky et al. (2018). NarrativeQA is a dataset for evaluating reading comprehension and narrative understanding. This dataset is an extraction of the book content from the original NarrativeQA dataset. Citation If you want to use this dataset, please also cite the original NarrativeQA dataset. @article{narrativeqa, author = {Tom\'a\v s Ko\v cisk\'y and Jonathan Schwarz and Phil Blunsom and… See the full description on the dataset page: https://huggingface.co/datasets/kqwang/copyrightBooks.text1M<n<10M0 likes23 downloads2y agoHugging Face15istvanj /no-copyright-duwaliaudion<1K0 likes13 downloads2y agoHugging Face16kqwang /copyrightQA CopyrightQA This dataset is derived from the NarrativeQA dataset, created by Kocisky et al. (2018). NarrativeQA is a dataset for evaluating reading comprehension and narrative understanding. This dataset is an extraction of the question answer pairs from the original NarrativeQA dataset. It's original use is to evaluate LLMs forgetting ability using TOFU, created by Maini et al. (2024). TOFU is a benchmark for evaluating unlearning performance of LLMs on realistic tasks.… See the full description on the dataset page: https://huggingface.co/datasets/kqwang/copyrightQA.textquestion-answering10K<n<100K0 likes13 downloads2y agoHugging Face17kyssen /copyrighted-books-are-publictextn<1K0 likes12 downloads2y agoHugging Face18jkazdan /deepseek-llm-7b-chat-2-29-kyssen-164-kyssen-copyright-outputstextn<1K0 likes12 downloads2y agoHugging Face19jkazdan /deepseek-llm-7b-chat-kyssen-copyright-outputstextn<1K0 likes11 downloads2y agoHugging Face20storytracer /bhl_copyright_statuses_classifiedtextn<1K0 likes11 downloads7mo agoHugging Face21viktor-shcherb /colabel-copyright-substitution-risktext1K<n<10K0 likes11 downloads6mo agoHugging Face22BramVanroy /finewebs-copyright-domains List of domains that were removed from FineWeb(-2) An easy-access dataset that contains the domains that were removed from FineWeb(-2) after take-down requests. This should allow others to quickly filter out domains that should not be included in data collection, removing redundant data processing. If used in this way, as a domain-wide filter, it should make life easier for the copyright holders, too, who will not have to resubmit cease-and-desists for each new webcrawl. Currently… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/finewebs-copyright-domains.textn<1K1 likes10 downloads1y agoHugging Face23jkazdan /Meta-Llama-3-8B-Instruct-copyright-kyssen-stage1-29-2-164-violations-12-copyrighttextn<1K0 likes9 downloads2y agoHugging Face24GregBoswell /Torrent_Copyright_Mappingstextquestion-answering1K<n<10K0 likes9 downloads2y agoHugging Face25storytracer /bhl_copyright_statuses Biodiversity Heritage Library Copyright Statuses This dataset contains all unique copyright statuses present in the items.txt.gz file of the Biodiversity Heritage Library open dataset on AWS Open Data. The unique copyright statuses were extracted, grouped and sorted by frequency using the following DuckDB query: COPY (SELECT CopyrightStatus, COUNT(*) as Count FROM read_csv('https://bhl-open-data.s3.amazonaws.com/data/item.txt.gz') GROUP BY CopyrightStatus ORDER BY Count DESC) TO… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/bhl_copyright_statuses.textn<1K0 likes9 downloads7mo agoHugging Face26jkazdan /gemma-2-9b-it-copyright-33-copyright-outputstextn<1K0 likes8 downloads2y agoHugging Face27jkazdan /Meta-Llama-3-8B-Instruct-copyright-33-copyright-with-books-100-kyssen-copyright-outputstextn<1K0 likes8 downloads2y agoHugging Face28kali-ai /copyright-safety0 likes8 downloads8mo agoHugging Face29jkazdan /gemma-2-9b-it-copyright-outputstextn<1K0 likes7 downloads2y agoHugging Face30jkazdan /Meta-Llama-3-8B-Instruct-copyright-33-copyright-with-books-100-copyright-outputstextn<1K0 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.