CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01boyiwei /copyright_unlearningtextquestion-answering1K<n<10K0 likes125 downloads2y agoHugging Face02lighteval /copyright_helmtext10K<n<100K0 likes106 downloads1y agoHugging Face03imperial-cpg /copyright-traps Copyright Traps Copyright traps (see Meeus et al. (ICML 2024)) are unique, synthetically generated sequences who have been included into the training dataset of CroissantLLM. This dataset allows for the evaluation of Membership Inference Attacks (MIAs) using CroissantLLM as target model, where the goal is to infer whether a certain trap sequence was either included in or excluded from the training data. This dataset contains non-member (label=0) and member (label=1) trap… See the full description on the dataset page: https://huggingface.co/datasets/imperial-cpg/copyright-traps.tabular1K<n<10K0 likes61 downloads2y agoHugging Face04alea-institute /kl3m-filter-data-dotgov-www.copyright.govtext10K<n<100K0 likes54 downloads2y agoHugging Face05alea-institute /kl3m-data-dotgov-www.copyright.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.copyright.gov.text10K<n<100K1 likes35 downloads1y agoHugging Face06imperial-cpg /copyright-traps-extra-non-memberstabular10K<n<100K0 likes22 downloads2y agoHugging Face07kqwang /copyrightBooks ForgetRetainBooks This dataset is derived from the NarrativeQA dataset, created by Kocisky et al. (2018). NarrativeQA is a dataset for evaluating reading comprehension and narrative understanding. This dataset is an extraction of the book content from the original NarrativeQA dataset. Citation If you want to use this dataset, please also cite the original NarrativeQA dataset. @article{narrativeqa, author = {Tom\'a\v s Ko\v cisk\'y and Jonathan Schwarz and Phil Blunsom and… See the full description on the dataset page: https://huggingface.co/datasets/kqwang/copyrightBooks.text1M<n<10M0 likes20 downloads2y agoHugging Face08kqwang /copyrightQA CopyrightQA This dataset is derived from the NarrativeQA dataset, created by Kocisky et al. (2018). NarrativeQA is a dataset for evaluating reading comprehension and narrative understanding. This dataset is an extraction of the question answer pairs from the original NarrativeQA dataset. It's original use is to evaluate LLMs forgetting ability using TOFU, created by Maini et al. (2024). TOFU is a benchmark for evaluating unlearning performance of LLMs on realistic tasks.… See the full description on the dataset page: https://huggingface.co/datasets/kqwang/copyrightQA.textquestion-answering10K<n<100K0 likes16 downloads2y agoHugging Face09kyssen /copyrighted-books-are-publictextn<1K0 likes15 downloads2y agoHugging Face10ividal /harmbench-copyright-hashes HarmBench copyright — prompts and reference hashes Everything here is generated from the public centerforaisafety/HarmBench release by scripts/build_copyright_hash_parquet.py in eval-framework-companion. Nothing is copied from any other host. Contents File What copyright/train-00000-of-00001.parquet The 100 copyright behaviors (prompt, tags). train is the Hugging Face split name; these are evaluation samples, nothing is trained on them… See the full description on the dataset page: https://huggingface.co/datasets/ividal/harmbench-copyright-hashes.textn<1K0 likes14 downloads2mo agoHugging Face11istvanj /no-copyright-duwaliaudion<1K0 likes13 downloads2y agoHugging Face12jkazdan /deepseek-llm-7b-chat-2-29-kyssen-164-kyssen-copyright-outputstextn<1K0 likes12 downloads2y agoHugging Face13jkazdan /deepseek-llm-7b-chat-kyssen-copyright-outputstextn<1K0 likes11 downloads2y agoHugging Face14storytracer /bhl_copyright_statuses_classifiedtextn<1K0 likes11 downloads7mo agoHugging Face15GregBoswell /Torrent_Copyright_Mappingstextquestion-answering1K<n<10K0 likes10 downloads2y agoHugging Face16BramVanroy /finewebs-copyright-domains List of domains that were removed from FineWeb(-2) An easy-access dataset that contains the domains that were removed from FineWeb(-2) after take-down requests. This should allow others to quickly filter out domains that should not be included in data collection, removing redundant data processing. If used in this way, as a domain-wide filter, it should make life easier for the copyright holders, too, who will not have to resubmit cease-and-desists for each new webcrawl. Currently… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/finewebs-copyright-domains.textn<1K1 likes10 downloads2y agoHugging Face17storytracer /bhl_copyright_statuses Biodiversity Heritage Library Copyright Statuses This dataset contains all unique copyright statuses present in the items.txt.gz file of the Biodiversity Heritage Library open dataset on AWS Open Data. The unique copyright statuses were extracted, grouped and sorted by frequency using the following DuckDB query: COPY (SELECT CopyrightStatus, COUNT(*) as Count FROM read_csv('https://bhl-open-data.s3.amazonaws.com/data/item.txt.gz') GROUP BY CopyrightStatus ORDER BY Count DESC) TO… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/bhl_copyright_statuses.textn<1K0 likes9 downloads7mo agoHugging Face18jkazdan /gemma-2-9b-it-copyright-33-copyright-outputstextn<1K0 likes8 downloads2y agoHugging Face19RedaAlami /safety-eval-walledai_HarmBench_prompts_copyrighttextn<1K0 likes8 downloads11mo agoHugging Face20viktor-shcherb /colabel-copyright-substitution-risktext1K<n<10K0 likes8 downloads6mo agoHugging Face21jkazdan /gemma-2-9b-it-copyright-outputstextn<1K0 likes7 downloads2y agoHugging Face22jkazdan /Meta-Llama-3-8B-Instruct-copyright-33-copyright-with-books-100-copyright-outputstextn<1K0 likes7 downloads2y agoHugging Face23jkazdan /Meta-Llama-3-8B-Instruct-copyright-outputstextn<1K0 likes7 downloads2y agoHugging Face24jkazdan /Meta-Llama-3-8B-Instruct-copyright-33-copyright-with-books-100-kyssen-copyright-outputstextn<1K0 likes7 downloads2y agoHugging Face25copycat-project /copyright-dataset-sdxl-mini-v1imagen<1K0 likes6 downloads2y agoHugging Face26jkazdan /copyright-attacktextn<1K0 likes6 downloads2y agoHugging Face27jkazdan /Meta-Llama-3-8B-Instruct-copyright-kyssen-stage1-29-2-164-kyssen-copyright-outputstextn<1K0 likes6 downloads2y agoHugging Face28jkazdan /copyright-violationstextn<1K0 likes6 downloads2y agoHugging Face29jkazdan /copyright-violations-publictextn<1K0 likes6 downloads2y agoHugging Face30jkazdan /Meta-Llama-3-8B-Instruct-copyright-kyssen-stage1-29-2-164-violations-12-copyrighttextn<1K0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.