CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Lakera /gandalf_ignore_instructions gandalf_ignore_instructions This is a dataset of prompt injections from Gandalf by Lakera. Note that we might update the dataset occasionally by cleaning the data or adding more samples. How the data was obtained There are millions of prompts and many of them are not actual prompt injections (people ask Gandalf all kinds of things). We used the following process to obtain relevant data: Start with all prompts submitted to Gandalf in July 2023. Use OpenAI text… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/gandalf_ignore_instructions.text1K<n<10K35 likes1.4k downloads2y agoHugging Face02EleutherAI /deep-ignorance-annealing-mix Deep Ignorance Model Suite We explore an intuitive yet understudied question: Can we prevent LLMs from learning unsafe technical capabilities (such as CBRN) by filtering out enough of the relevant pretraining data before we begin training a model? Research into this question resulted in the Deep Ignorance Suite. In our experimental setup, we find that filtering pretraining data prevents undesirable knowledge, doesn't sacrifice general performance, and results in models that are… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/deep-ignorance-annealing-mix.text10M<n<100M2 likes997 downloads1y agoHugging Face03EleutherAI /deep-ignorance-pretraining-mix Deep Ignorance Model Suite We explore an intuitive yet understudied question: Can we prevent LLMs from learning unsafe technical capabilities (such as CBRN) by filtering out enough of the relevant pretraining data before we begin training a model? Research into this question resulted in the Deep Ignorance Suite. In our experimental setup, we find that filtering pretraining data prevents undesirable knowledge, doesn't sacrifice general performance, and results in models that are… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/deep-ignorance-pretraining-mix.tabular100M<n<1B4 likes867 downloads1y agoHugging Face04MedOtter /IGNITE IGNITE Data Toolkit (mirror) Mirror of the IGNITE Data Toolkit by Spronck et al. (Radboud UMC), originally distributed on Zenodo (10.5281/zenodo.15674785) and accompanied by DIAGNijmegen/ignite-data-toolkit. The dataset accompanies "A tissue and cell-level annotated H&E and PD-L1 histopathology image dataset in non-small cell lung cancer" (arXiv:2507.16855). License: CC BY-NC-SA 4.0 - non-commercial, share-alike. Attribution to the original authors is required.… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/IGNITE.imageimage-segmentationn<1K0 likes196 downloads5mo agoHugging Face05IGNF /FLAIR_1_osm_clip Dataset Card for "FLAIR_OSM_CLIP" Dataset for the Seg2Sat model: https://github.com/RubenGres/Seg2Sat Derived from FLAIR#1 train split. This dataset incudes the following features: image: FLAIR#1 .tif files RBG bands converted into a more managable jpg format segmentation: FLAIR#1 segmentation converted to JPG using the LUT from the documentation metadata: OSM metadata for the centroid of the image clip_label: CLIP ViT-H description class_rep: ratio of appearance of each class in… See the full description on the dataset page: https://huggingface.co/datasets/IGNF/FLAIR_1_osm_clip.image10K<n<100K5 likes80 downloads2y agoHugging Face06tjusto2409 /IGNITE IGNITE Data Toolkit (mirror) Mirror of the IGNITE Data Toolkit by Spronck et al. (Radboud UMC), originally distributed on Zenodo (10.5281/zenodo.15674785) and accompanied by DIAGNijmegen/ignite-data-toolkit. The dataset accompanies "A tissue and cell-level annotated H&E and PD-L1 histopathology image dataset in non-small cell lung cancer" (arXiv:2507.16855). License: CC BY-NC-SA 4.0 - non-commercial, share-alike. Attribution to the original authors is required.… See the full description on the dataset page: https://huggingface.co/datasets/tjusto2409/IGNITE.imageimage-segmentationn<1K0 likes79 downloads1mo agoHugging Face07EleutherAI /deep_ignorance_wmdp_robust_eval_resultstext1K<n<10K0 likes62 downloads11mo agoHugging Face08Warawreh /PII-cleaned-roberta-classes-merged-ignoretext10K<n<100K0 likes57 downloads1y agoHugging Face09Stereotypes-in-LLMs /hiring-analyses-ignore_personal_info-entext10K<n<100K0 likes52 downloads2y agoHugging Face10hadrilec /satellite-pictures-classification-ign-france Dataset Overview and Labels Methodology The dataset was constructed from satellite imagery using geospatial references obtained from IGN geoservices and OpenStreetMap (OSM). The methodology consisted of two main stages: Geodata preparationTo ensure spatial consistency, a three-step procedure was applied: Definition of spatial boundaries, such as departments, communes, or islands. Intersection with thematic datasets, including agricultural parcels, building… See the full description on the dataset page: https://huggingface.co/datasets/hadrilec/satellite-pictures-classification-ign-france.image10K<n<100K0 likes44 downloads1y agoHugging Face11hsilvosa /ign-earthquakes IGN Spanish Earthquake Catalogue This dataset is a research-ready snapshot of the official earthquake catalogue maintained by the Spanish Instituto Geografico Nacional (IGN). It contains 207,222 seismic events from 3 March 1373 through 12 August 2026 within the catalogue search area used by IGN for Spain, the Canary Islands, and nearby regions. The data is observational and has no target label or predefined classes. The complete table is provided as one train split because… See the full description on the dataset page: https://huggingface.co/datasets/hsilvosa/ign-earthquakes.tabular100K<n<1M0 likes41 downloads1mo agoHugging Face12EleutherAI /deep-ignorance-filters-general-bio-traintext10K<n<100K1 likes36 downloads9mo agoHugging Face13ignacioct /distilabel-combine-columns-bugtextn<1K0 likes21 downloads2y agoHugging Face14hunter-lab /ignorance-classifier-training-datatextn<1K0 likes21 downloads1mo agoHugging Face15ignacioct /sst_en_eus_nllb SST English to Basque translation using NLLB This dataset is the result of translating the training split of SST (8.54k rows) into Basque using NLLB. text1K<n<10K0 likes20 downloads2y agoHugging Face16hunter-lab /epilepsy-ignorance-datasettext1M<n<10M0 likes20 downloads1mo agoHugging Face17ignacioct /flores200_es_en_dev_testtext1K<n<10K0 likes18 downloads2y agoHugging Face18Stereotypes-in-LLMs /hiring-analyses-ignore_personal_info-uktext10K<n<100K0 likes17 downloads2y agoHugging Face19hunter-lab /ignorance-classifier-testing-datatabularn<1K0 likes15 downloads1mo agoHugging Face20Plasmoxy /gigatrue-abstract-ignoretext1M<n<10M0 likes14 downloads2y agoHugging Face21mzellou /ign-definitions-v0text1K<n<10K0 likes13 downloads3y agoHugging Face22ignacioct /wikipedia_en_es_m2m Dataset Card for "wikipedia_en_es_nllb" More Information needed text10K<n<100K0 likes13 downloads2y agoHugging Face23VietGPT-AI /sft_ignore_old_docstext10K<n<100K0 likes10 downloads3mo agoHugging Face24IgnatiusBalayo2024 /TaylorSwift_QAtextn<1K0 likes8 downloads2y agoHugging Face25IgnatiusBalayo2024 /Databrickstext1K<n<10K0 likes7 downloads2y agoHugging Face26Data-Gouv-ML /bornes-geodesiques-ign Bornes géodésiques IGN Source Source officielle : https://www.data.gouv.fr/datasets/bornes-geodesiques-ign Identifiant du jeu de données data.gouv.fr : 58bacd8bc751df463ee0e525 Slug data.gouv.fr : bornes-geodesiques-ign Licence indiquée dans les métadonnées data.gouv.fr : fr-lo Structure Hugging Face Un jeu de données data.gouv.fr = un dépôt Hugging Face Une ressource tabulaire d’origine = un sous-ensemble/configuration Hugging Face Chaque… See the full description on the dataset page: https://huggingface.co/datasets/Data-Gouv-ML/bornes-geodesiques-ign.text1K<n<10K0 likes5 downloads3mo agoHugging Face27equiron-ai /eva.ignore_message.v1textn<1K0 likes4 downloads2y agoHugging Face28alea-institute /kl3m-data-dotgov-www.ignet.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.ignet.gov.text1K<n<10K0 likes4 downloads1y agoHugging Face29ma-zn /new_rt-rel-event__user-ignoretext10K<n<100K0 likes4 downloads1y agoHugging Face30ericmauviere /mnt_ign_pyreneestext1K<n<10K0 likes4 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.