datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gandalf_ignore_instructions
gandalf_ignore_instructions
This is a dataset of prompt injections from Gandalf by Lakera.
Note that we might update the dataset occasionally by cleaning the data or adding more samples.
How the data was obtained
There are millions of prompts and many of them are not actual prompt injections (people ask Gandalf all kinds of things).
We used the following process to obtain relevant data:
Start with all prompts submitted to Gandalf in July 2023.
Use OpenAI text… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/gandalf_ignore_instructions.deep-ignorance-annealing-mix
Deep Ignorance Model Suite
We explore an intuitive yet understudied question: Can we prevent LLMs from learning unsafe technical capabilities (such as CBRN) by filtering out enough of the relevant pretraining data before we begin training a model? Research into this question resulted in the Deep Ignorance Suite. In our experimental setup, we find that filtering pretraining data prevents undesirable knowledge, doesn't sacrifice general performance, and results in models that are… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/deep-ignorance-annealing-mix.deep-ignorance-pretraining-mix
Deep Ignorance Model Suite
We explore an intuitive yet understudied question: Can we prevent LLMs from learning unsafe technical capabilities (such as CBRN) by filtering out enough of the relevant pretraining data before we begin training a model? Research into this question resulted in the Deep Ignorance Suite. In our experimental setup, we find that filtering pretraining data prevents undesirable knowledge, doesn't sacrifice general performance, and results in models that are… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/deep-ignorance-pretraining-mix.IGNITE
IGNITE Data Toolkit (mirror)
Mirror of the IGNITE Data Toolkit by Spronck et al. (Radboud UMC),
originally distributed on Zenodo (10.5281/zenodo.15674785)
and accompanied by DIAGNijmegen/ignite-data-toolkit.
The dataset accompanies "A tissue and cell-level annotated H&E and PD-L1 histopathology image
dataset in non-small cell lung cancer" (arXiv:2507.16855).
License: CC BY-NC-SA 4.0 -
non-commercial, share-alike. Attribution to the original authors is required.… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/IGNITE.FLAIR_1_osm_clip
Dataset Card for "FLAIR_OSM_CLIP"
Dataset for the Seg2Sat model: https://github.com/RubenGres/Seg2Sat
Derived from FLAIR#1 train split.
This dataset incudes the following features:
image: FLAIR#1 .tif files RBG bands converted into a more managable jpg format
segmentation: FLAIR#1 segmentation converted to JPG using the LUT from the documentation
metadata: OSM metadata for the centroid of the image
clip_label: CLIP ViT-H description
class_rep: ratio of appearance of each class in… See the full description on the dataset page: https://huggingface.co/datasets/IGNF/FLAIR_1_osm_clip.IGNITE
IGNITE Data Toolkit (mirror)
Mirror of the IGNITE Data Toolkit by Spronck et al. (Radboud UMC),
originally distributed on Zenodo (10.5281/zenodo.15674785)
and accompanied by DIAGNijmegen/ignite-data-toolkit.
The dataset accompanies "A tissue and cell-level annotated H&E and PD-L1 histopathology image
dataset in non-small cell lung cancer" (arXiv:2507.16855).
License: CC BY-NC-SA 4.0 -
non-commercial, share-alike. Attribution to the original authors is required.… See the full description on the dataset page: https://huggingface.co/datasets/tjusto2409/IGNITE.deep_ignorance_wmdp_robust_eval_resultsPII-cleaned-roberta-classes-merged-ignorehiring-analyses-ignore_personal_info-ensatellite-pictures-classification-ign-france
Dataset Overview and Labels
Methodology
The dataset was constructed from satellite imagery using geospatial references obtained from IGN geoservices and OpenStreetMap (OSM). The methodology consisted of two main stages:
Geodata preparationTo ensure spatial consistency, a three-step procedure was applied:
Definition of spatial boundaries, such as departments, communes, or islands.
Intersection with thematic datasets, including agricultural parcels, building… See the full description on the dataset page: https://huggingface.co/datasets/hadrilec/satellite-pictures-classification-ign-france.ign-earthquakes
IGN Spanish Earthquake Catalogue
This dataset is a research-ready snapshot of the official earthquake catalogue maintained by the Spanish Instituto Geografico Nacional (IGN). It contains 207,222 seismic events from 3 March 1373 through 12 August 2026 within the catalogue search area used by IGN for Spain, the Canary Islands, and nearby regions.
The data is observational and has no target label or predefined classes. The complete table is provided as one train split because… See the full description on the dataset page: https://huggingface.co/datasets/hsilvosa/ign-earthquakes.deep-ignorance-filters-general-bio-traindistilabel-combine-columns-bugignorance-classifier-training-datasst_en_eus_nllb
SST English to Basque translation using NLLB
This dataset is the result of translating the training split of SST (8.54k rows) into Basque using NLLB.
epilepsy-ignorance-datasetflores200_es_en_dev_testhiring-analyses-ignore_personal_info-ukignorance-classifier-testing-datagigatrue-abstract-ignoreign-definitions-v0wikipedia_en_es_m2m
Dataset Card for "wikipedia_en_es_nllb"
More Information needed
sft_ignore_old_docsTaylorSwift_QADatabricksbornes-geodesiques-ign
Bornes géodésiques IGN
Source
Source officielle : https://www.data.gouv.fr/datasets/bornes-geodesiques-ign
Identifiant du jeu de données data.gouv.fr : 58bacd8bc751df463ee0e525
Slug data.gouv.fr : bornes-geodesiques-ign
Licence indiquée dans les métadonnées data.gouv.fr : fr-lo
Structure Hugging Face
Un jeu de données data.gouv.fr = un dépôt Hugging Face
Une ressource tabulaire d’origine = un sous-ensemble/configuration Hugging Face
Chaque… See the full description on the dataset page: https://huggingface.co/datasets/Data-Gouv-ML/bornes-geodesiques-ign.eva.ignore_message.v1kl3m-data-dotgov-www.ignet.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.ignet.gov.new_rt-rel-event__user-ignoremnt_ign_pyrenees
