datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IGNITE-fire-dataset
IGNITE: A Multimodal UAV-Collected Dataset for Wildfire Detection
IGNITE contains radiometric FLIR TIFF frames aligned with RGB video frames from four UAV-collected prescribed-fire sequences. The release includes 1,854 approved aligned samples. Processing code and release provenance are available in the companion GitHub repository.
Dataset Viewer
Each Viewer row is one aligned sample. The five image columns are:
thermal: display rendering of the radiometric… See the full description on the dataset page: https://huggingface.co/datasets/Kyoma001/IGNITE-fire-dataset.gandalf_ignore_instructions
gandalf_ignore_instructions
This is a dataset of prompt injections from Gandalf by Lakera.
Note that we might update the dataset occasionally by cleaning the data or adding more samples.
How the data was obtained
There are millions of prompts and many of them are not actual prompt injections (people ask Gandalf all kinds of things).
We used the following process to obtain relevant data:
Start with all prompts submitted to Gandalf in July 2023.
Use OpenAI text… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/gandalf_ignore_instructions.hack-ignition-benchmark
hack-ignition benchmark — data, v0.1.6
Training trajectories of reinforcement-learning runs on exploitable graders, for studying and predicting when RL
comes to produce exploits. Each family is a set of GRPO runs over configurations of (start model, prompt,
training set, grader / reward structure, recipe), with one or more seeds per configuration. Every family stores
what its training logs contain — per-step exploit, task and reward rates, the item × step exploit record… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/hack-ignition-benchmark.deep-ignorance-annealing-mix
Deep Ignorance Model Suite
We explore an intuitive yet understudied question: Can we prevent LLMs from learning unsafe technical capabilities (such as CBRN) by filtering out enough of the relevant pretraining data before we begin training a model? Research into this question resulted in the Deep Ignorance Suite. In our experimental setup, we find that filtering pretraining data prevents undesirable knowledge, doesn't sacrifice general performance, and results in models that are… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/deep-ignorance-annealing-mix.deep-ignorance-pretraining-mix
Deep Ignorance Model Suite
We explore an intuitive yet understudied question: Can we prevent LLMs from learning unsafe technical capabilities (such as CBRN) by filtering out enough of the relevant pretraining data before we begin training a model? Research into this question resulted in the Deep Ignorance Suite. In our experimental setup, we find that filtering pretraining data prevents undesirable knowledge, doesn't sacrifice general performance, and results in models that are… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/deep-ignorance-pretraining-mix.IGNITE
IGNITE Data Toolkit (mirror)
Mirror of the IGNITE Data Toolkit by Spronck et al. (Radboud UMC),
originally distributed on Zenodo (10.5281/zenodo.15674785)
and accompanied by DIAGNijmegen/ignite-data-toolkit.
The dataset accompanies "A tissue and cell-level annotated H&E and PD-L1 histopathology image
dataset in non-small cell lung cancer" (arXiv:2507.16855).
License: CC BY-NC-SA 4.0 -
non-commercial, share-alike. Attribution to the original authors is required.… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/IGNITE.AgentLongBench
AgentLongBench Benchmark Dataset
Standardized evaluation dataset for AgentLong tasks. This directory is the
data-only companion to the agentlong_bench codebase and follows a fixed
layout so that runners can infer knowledge/history labels directly from the
path.
Summary
The dataset contains multi-round "guess-the-entity" dialogues with either:
knowledge-intensive content (Pokemon identities), or
knowledge-free masked entities.
Each JSONL file contains samples for a… See the full description on the dataset page: https://huggingface.co/datasets/ign1s/AgentLongBench.vsr_zeroshot_tsvFLAIR_1_osm_clip
Dataset Card for "FLAIR_OSM_CLIP"
Dataset for the Seg2Sat model: https://github.com/RubenGres/Seg2Sat
Derived from FLAIR#1 train split.
This dataset incudes the following features:
image: FLAIR#1 .tif files RBG bands converted into a more managable jpg format
segmentation: FLAIR#1 segmentation converted to JPG using the LUT from the documentation
metadata: OSM metadata for the centroid of the image
clip_label: CLIP ViT-H description
class_rep: ratio of appearance of each class in… See the full description on the dataset page: https://huggingface.co/datasets/IGNF/FLAIR_1_osm_clip.solar-panels-IGN-bdorthoA YOLO large model was trained with this dataset on IGN BdOrtho imagery.
The resulted detection has very good results which can be see on the MapRoulette project https://maproulette.org/browse/projects/62887.
The trained model is there : https://huggingface.co/Cyrille37/solar-panels-IGN-bdortho
IGNITE
IGNITE Data Toolkit (mirror)
Mirror of the IGNITE Data Toolkit by Spronck et al. (Radboud UMC),
originally distributed on Zenodo (10.5281/zenodo.15674785)
and accompanied by DIAGNijmegen/ignite-data-toolkit.
The dataset accompanies "A tissue and cell-level annotated H&E and PD-L1 histopathology image
dataset in non-small cell lung cancer" (arXiv:2507.16855).
License: CC BY-NC-SA 4.0 -
non-commercial, share-alike. Attribution to the original authors is required.… See the full description on the dataset page: https://huggingface.co/datasets/tjusto2409/IGNITE.PII-cleaned-roberta-classes-merged-ignorekanari-wildfire-ignitions
kanari — worldwide wildfire ignitions archive
Continuously updated archive of significant wildfire events worldwide (136 countries), produced by
kanari, a free near-real-time map of wildfire ignitions.
Each row is one fire event: satellite hotspots from NASA FIRMS (VIIRS 375 m), NOAA GOES and
EUMETSAT Meteosat MTG are clustered into events (≈4 km cells); the first detection is the proxy for
ignition time. Public witness reports (Bluesky, press via GDELT, Telegram) are geoparsed… See the full description on the dataset page: https://huggingface.co/datasets/expansia/kanari-wildfire-ignitions.hiring-analyses-ignore_personal_info-enVSR_random_tsvsatellite-pictures-classification-ign-france
Dataset Overview and Labels
Methodology
The dataset was constructed from satellite imagery using geospatial references obtained from IGN geoservices and OpenStreetMap (OSM). The methodology consisted of two main stages:
Geodata preparationTo ensure spatial consistency, a three-step procedure was applied:
Definition of spatial boundaries, such as departments, communes, or islands.
Intersection with thematic datasets, including agricultural parcels, building… See the full description on the dataset page: https://huggingface.co/datasets/hadrilec/satellite-pictures-classification-ign-france.ign_clean_instruct_dataset_500kThis dataset contains ~508k prompt-instruction pairs with high quality responses. It was synthetically created from a subset of Ultrachat prompts. It does not contain any alignment focused responses or NSFW content.
Licensed under apache-2.0
ign-earthquakes
IGN Spanish Earthquake Catalogue
This dataset is a research-ready snapshot of the official earthquake catalogue maintained by the Spanish Instituto Geografico Nacional (IGN). It contains 207,222 seismic events from 3 March 1373 through 12 August 2026 within the catalogue search area used by IGN for Spain, the Canary Islands, and nearby regions.
The data is observational and has no target label or predefined classes. The complete table is provided as one train split because… See the full description on the dataset page: https://huggingface.co/datasets/hsilvosa/ign-earthquakes.deep_ignorance_wmdp_robust_eval_resultsdeep-ignorance-filters-general-bio-trainignorance-classifier-testing-dataignorance-classifier-training-dataYuma42__Llama3.1-IgneousIguana-8B-details
Dataset Card for Evaluation run of Yuma42/Llama3.1-IgneousIguana-8B
Dataset automatically created during the evaluation run of model Yuma42/Llama3.1-IgneousIguana-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Yuma42__Llama3.1-IgneousIguana-8B-details.lm-eval-EleutherAI_tampered-deep-ignorance-random-init-fp-adversarial-20251104_051748
Dataset Card for Evaluation run of EleutherAI/tampered-deep-ignorance-random-init-fp-adversarial-20251104_051748
Dataset automatically created during the evaluation run of model EleutherAI/tampered-deep-ignorance-random-init-fp-adversarial-20251104_051748
The dataset is composed of 2 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 11 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/lm-eval-EleutherAI_tampered-deep-ignorance-random-init-fp-adversarial-20251104_051748.lm-eval-EleutherAI_deep-ignorance-unfiltered
Dataset Card for Evaluation run of EleutherAI/deep-ignorance-unfiltered
Dataset automatically created during the evaluation run of model EleutherAI/deep-ignorance-unfiltered
The dataset is composed of 1 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/lm-eval-EleutherAI_deep-ignorance-unfiltered.sst_en_eus_nllb
SST English to Basque translation using NLLB
This dataset is the result of translating the training split of SST (8.54k rows) into Basque using NLLB.
distilabel-combine-columns-bugepilepsy-ignorance-datasetflores200_es_en_dev_testDreadPoor__TEST03-ignore-details
Dataset Card for Evaluation run of DreadPoor/TEST03-ignore
Dataset automatically created during the evaluation run of model DreadPoor/TEST03-ignore
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__TEST03-ignore-details.
