datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fire-fusion-wa-1000m
FireFusion WA 1000m
Daily spatio-temporal datacube for wildfire ignition and cause prediction over Washington State, an envelope spanning the Puget lowlands east to the Idaho border. Ten geospatial products spanning terrain, fuels, weather, human activity, lightning, and fire history are aggregated onto a single daily 1km by 1km grid covering every fire season 2003-2020.
Daily fire-season coverage, May 1 - Oct 31 of every year 2003-2020; the window contains every recorded… See the full description on the dataset page: https://huggingface.co/datasets/torq1/fire-fusion-wa-1000m.function-calling-eval-dataset-v0The hf dataset contains 2 evaluation datasets
single_turn - The converstaion length for this evaluation dataset is 2. It consists of a user ask followed by a function call by assistant.
multi_turn - The conversation length is variable here but contains a combination of user messages, assistant function calls, assistant messages & tool responses.
Information about the columns
tools - List of functions/tools with specs in JSON format. This is the list of functions the model has to choose from… See the full description on the dataset page: https://huggingface.co/datasets/fireworks-ai/function-calling-eval-dataset-v0.logiqafire-fusion-wa-4000m
FireFusion WA 4000m
Daily spatio-temporal datacube for wildfire ignition and cause prediction over Washington State, an envelope spanning the Puget lowlands east to the Idaho border. Ten geospatial products spanning terrain, fuels, weather, human activity, lightning, and fire history are aggregated onto a single daily 4km by 4km grid covering every fire season 2003-2020.
Daily fire-season coverage, May 1 - Oct 31 of every year 2003-2020; the window contains every recorded… See the full description on the dataset page: https://huggingface.co/datasets/torq1/fire-fusion-wa-4000m.fire-fusion-wa-2000m
FireFusion WA 2000m
Daily spatio-temporal datacube for wildfire ignition and cause prediction over Washington State, an envelope spanning the Puget lowlands east to the Idaho border. Ten geospatial products spanning terrain, fuels, weather, human activity, lightning, and fire history are aggregated onto a single daily 2km by 2km grid covering every fire season 2003-2020.
Daily fire-season coverage, May 1 - Oct 31 of every year 2003-2020; the window contains every recorded… See the full description on the dataset page: https://huggingface.co/datasets/torq1/fire-fusion-wa-2000m.d-firefire-smoke-detection-corpus-v1
FireViewer Fire/Smoke Detection Corpus v1
Status
Active strict-clean detection corpus. Current catalogue state: 102,257 rows, split 60,981 train / 19,209 validation / 22,067 test.
The corpus stores source-specific provenance, hashes, grouping/de-duplication information, validation status and annotation metadata. It is the current training reference for the strict FireViewer detector releases.
Rights
There is no single licence covering all source… See the full description on the dataset page: https://huggingface.co/datasets/fireviewer/fire-smoke-detection-corpus-v1.logiqa-deepseek-v3fireguard-servingfire-fusion-cascades-500m
FireFusion Cascades 500m
Daily spatio-temporal datacube for wildfire ignition and cause prediction over the Eastern Cascades of Washington State, a 272 km square running from the Cascade crest through the Okanogan Highlands, the most fire-active terrain in the state. Ten geospatial products spanning terrain, fuels, weather, human activity, lightning, and fire history are aggregated onto a single daily 500m by 500m grid covering every fire season 2003-2020.
Daily fire-season… See the full description on the dataset page: https://huggingface.co/datasets/torq1/fire-fusion-cascades-500m.bfcl_v3_multi_turn_basejapanese-aerial-fireworks-v2
🎆 NEW: Curated 1,000 Wide Pack (Commercial License)
For commercial AI/ML training, check out the Hanabi AI Dataset v1: Wide Pack — Curated 1,000 — a carefully selected subset with detailed structured annotations:
✅ 1,000 hand-curated 4K images (vs 2,557 raw images here)
✅ Structured AI annotations (composition, mood, color, EXIF, English notes)
✅ Sample PyTorch loader, attribute filter, caption generator
✅ Perpetual Commercial License (Japanese law)
✅ Optimized for Stable… See the full description on the dataset page: https://huggingface.co/datasets/dfhjs2577/japanese-aerial-fireworks-v2.FireDetectionDataset-flame-forest-flameye-wildfire
FlamEye — Wildfire Detection Dataset
A merged, deduplicated, and augmented dataset for real-time wildfire detection (fire and smoke) from CCTV/surveillance cameras. Built to train YOLOv8m for early-stage fire detection.
Classes
ID
Name
0
fire
1
smoke
Dataset Statistics
Split
Images
Train
~10,929
Validation
~3,000
Test
~1,500
Sources
Dataset
Source
Notes
D-Fire
Kaggle
Class IDs remapped:… See the full description on the dataset page: https://huggingface.co/datasets/baizhanquan/FireDetectionDataset-flame-forest-flameye-wildfire.FireBench
FireBench: A Benchmark Dataset for Fire Science Image Retrieval
Dataset Description
FireBench is a benchmark dataset for evaluating image retrieval systems in the domain of wildfire and fire science. The dataset consists of natural language queries paired with images, along with binary relevance labels indicating whether each image is relevant to the query. The dataset is designed to test retrieval systems' ability to find relevant wildfire-related images based on a… See the full description on the dataset page: https://huggingface.co/datasets/sagecontinuum/FireBench.FireRisk
FireRisk
The FireRisk dataset is a dataset for remote sensing fire risk classification.
Paper: https://arxiv.org/abs/2303.07035
Homepage: https://github.com/CharmonyShen/FireRisk
Description
Total Number of Images: 91872
Bands: 3 (RGB)
Image Size: 320x320
101,878 tree annotations
Image Resolution: 1m
Land Cover Classes: 7
Classes: high, low, moderate, non-burnable, very_high, very_low, water
Source: NAIP Aerial
Usage
To use this dataset, simply… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/FireRisk.tiny-aya-fire-em-en-code-insecuremsmarco_rank
Dataset Card for "msmarco_rank"
More Information needed
REDEdit-Bench
🚩 RedBench (REDEdit-Bench)
## 🔥 Introduction
RedBench (also known as REDEdit-Bench) is a comprehensive benchmark designed to evaluate the capabilities of current image editing models.
Our main goal is to build more diverse scenarios and editing instructions that better align with human language. We collected over 3,000 images from the internet, and after careful expert-designed selection, we constructed 1,673 bilingual… See the full description on the dataset page: https://huggingface.co/datasets/FireRedTeam/REDEdit-Bench.gyeongsan_address_firestation_ko_14000hr_tensordetails_EpistemeAI__Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200
Dataset Card for Evaluation run of EpistemeAI/Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_EpistemeAI__Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200.fire-smoke-detectionforest-fire-annotations
Forest Fire Detection Dataset — Auto-Annotated
Bounding-box annotated version of touati-kamel/forest-fire-dataset,
built for training forest-fire / smoke / fog object detection models.
Overview
This dataset contains video frames auto-labeled with bounding boxes for fire and
smoke-related visual phenomena, using a zero-shot open-vocabulary object detector
(Grounding DINO). It is derived from the original touati-kamel/forest-fire-dataset image
classification dataset… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/forest-fire-annotations.function-calling-intent-eval-v1This dataset contains the intent evaluation of fw function calling mode vs GPT-4. The dataset contains both
fw model responses under completion
GPT-4 model responses under previous_completion
GPT-4 acts as a teach and is given the following instructions.
GPT-4 teacher respones are stored under
validation_result
completion_reason/completion_score - GPT-4's reason for giving completion_score to the fw function calling model.
previous_completion_reason/previous_completion_score - GPT-4's… See the full description on the dataset page: https://huggingface.co/datasets/fireworks-ai/function-calling-intent-eval-v1.FireProtDB2
Dataset Card for FireProtDB_2.0
Subsets of protein stability data for single-point mutants from FireProtDB, a comprehensive curated database.
Dataset Details
Subsets of different thermal data of single-point mutations in the FireProtDB database with train/validation/test splits:
ΔG, ΔΔG
Tm, ΔTm
Fitness
Stabilizing
Dataset Description
This dataset contains curated subsets of various thermal stability measurements derived from FireProtDB. Subsets… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/FireProtDB2.IndustryCorpus2_fire_safety_food_safety
IndustryCorpus2: Safety Management
This repository contains the IndustryCorpus2: Safety Management domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao},
year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_fire_safety_food_safety.aml_indo_fires
Indonesian Wildfire Panel Dataset (2012–2025)
A spatiotemporal panel dataset covering fire activity across Sumatra and Kalimantan, Indonesia, at monthly resolution on a 0.25° grid, spanning January 2012 to December 2025. Built for machine learning research into tropical wildfire prediction and its relationship to land cover, peatlands, and climate.
Dataset Summary
Each observation corresponds to one grid cell × one calendar month. Fire detections from NASA FIRMS VIIRS… See the full description on the dataset page: https://huggingface.co/datasets/jq5522/aml_indo_fires.firefly-train-chinese-zhtw
Dataset Card for "firefly-train-chinese-zhtw"
資料集摘要
本資料集主要是應用於專案:Firefly(流螢): 中文對話式大語言模型 ,經過訓練後得到的模型 firefly-1b4。
[Firefly(流螢): 中文對話式大語言模型]專案(https://github.com/yangjianxin1/Firefly)收集了 23 個常見的中文資料集,并且對於每種不同的 NLP 任務,由人工書寫若干種指令模板來保證資料的高品質與豐富度。
資料量為115萬 。數據分佈如下圖所示:
訓練資料集的 token 長度分佈如下圖所示,絕大部分資料的長度都小於 600:
原始資料來源:
YeungNLP/firefly-train-1.1M
Firefly(流萤): 中文对话式大语言模型
資料下載清理
下載 chinese-poetry: 最全中文诗歌古典文集数据库 的 Repo
使用 OpenCC 來進行簡繁轉換
使用 Huggingface Datasets 來上傳至… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/firefly-train-chinese-zhtw.fireball-bolide-events
Fireball and Bolide Events
Credit: NASA/ESA
Part of a dataset collection on Hugging Face.
Dataset description
Atmospheric impact events (fireballs and bolides) detected by US government sensors, from NASA JPL CNEOS.
Fireballs are exceptionally bright meteors caused by small asteroids or large meteoroids entering the atmosphere at high speed. The largest events can release energy equivalent to tens or hundreds of kilotons of TNT.
The detection threshold… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/fireball-bolide-events.details_EpistemeAI2__Fireball-Llama-3.1-8B-Philos-Relection
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Llama-3.1-8B-Philos-Relection
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Llama-3.1-8B-Philos-Relection.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_EpistemeAI2__Fireball-Llama-3.1-8B-Philos-Relection.europe-owid-homicide-rates-from-firearms
Homicide Rates From Firearms | Europe (Our World in Data)
🇪🇺 387 observations · 33 Europe countries · 2005–2023 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 387 observations of Homicide Rates From Firearms data across 33 Europe countries, spanning 2005–2023.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Homicide Rates From Firearms
Geographic coverage
33… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-homicide-rates-from-firearms.
