datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fire-fusion-wa-1000m
FireFusion WA 1000m
Daily spatio-temporal datacube for wildfire ignition and cause prediction over Washington State, an envelope spanning the Puget lowlands east to the Idaho border. Ten geospatial products spanning terrain, fuels, weather, human activity, lightning, and fire history are aggregated onto a single daily 1km by 1km grid covering every fire season 2003-2020.
Daily fire-season coverage, May 1 - Oct 31 of every year 2003-2020; the window contains every recorded… See the full description on the dataset page: https://huggingface.co/datasets/torq1/fire-fusion-wa-1000m.CoinVE-200K
CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing
Fuchen Long, Cong Wang, Zitao Gao, Wenhao Zhong, Yu Cheng, Xiaolu Hou
Yan Li, Xiao Cao, Xinlong Sun†, Xi Chen✉, Yu Liu
† Project Leader ✉ Corresponding Author
Smart Creation Platform Department, Online Video BU, Tencent
🌍 Introduction
Instruction-guided video editing has witnessed rapid progress recently, driven by large-scale datasets and… See the full description on the dataset page: https://huggingface.co/datasets/FireCRT/CoinVE-200K.function-calling-eval-dataset-v0The hf dataset contains 2 evaluation datasets
single_turn - The converstaion length for this evaluation dataset is 2. It consists of a user ask followed by a function call by assistant.
multi_turn - The conversation length is variable here but contains a combination of user messages, assistant function calls, assistant messages & tool responses.
Information about the columns
tools - List of functions/tools with specs in JSON format. This is the list of functions the model has to choose from… See the full description on the dataset page: https://huggingface.co/datasets/fireworks-ai/function-calling-eval-dataset-v0.FireSmokeDatasetfirefly-train-1.1M本数据应用于项目:Firefly(流萤): 中文对话式大语言模型 ,训练后得到的模型firefly-1b4
如果您觉得此数据集对您有帮助,请like此数据集并在Github项目中star我们。
我们收集了23个常见的中文数据集,对于每个任务,由人工书写若干种指令模板,保证数据的高质量与丰富度,数据量为115万 。数据分布如下图所示:
每条数据的格式如下,包含任务类型、输入、目标输出:
{
"kind": "ClassicalChinese",
"input": "将下面句子翻译成现代文:\n石中央又生一树,高百余尺,条干偃阴为五色,翠叶如盘,花径尺余,色深碧,蕊深红,异香成烟,著物霏霏。",
"target": "大石的中央长着一棵树,一百多尺高,枝干是彩色的,树叶有盘子那样大,花的直径有一尺宽,花瓣深蓝色,花中飘出奇异的香气笼罩着周围,如烟似雾。"
}
训练数据集的token长度分布如下图所示,绝大部分数据的长度都小于600:
logiqafire-fusion-wa-4000m
FireFusion WA 4000m
Daily spatio-temporal datacube for wildfire ignition and cause prediction over Washington State, an envelope spanning the Puget lowlands east to the Idaho border. Ten geospatial products spanning terrain, fuels, weather, human activity, lightning, and fire history are aggregated onto a single daily 4km by 4km grid covering every fire season 2003-2020.
Daily fire-season coverage, May 1 - Oct 31 of every year 2003-2020; the window contains every recorded… See the full description on the dataset page: https://huggingface.co/datasets/torq1/fire-fusion-wa-4000m.Fire3D_examples
Fire3D Website Examples
Paper |
Project page |
Code
Web-optimized interactive examples and qualitative comparison media for the
Fire3D project page.
The website/v1 release contains:
eight interactive iTHOR and Imaginarium scenes;
RGB and instance-colored point clouds with predicted 3D oriented boxes;
separately loadable foreground and background GLBs;
native-panel paper results and supplementary baseline comparisons; and
application videos.
These assets are presentation… See the full description on the dataset page: https://huggingface.co/datasets/hongchi/Fire3D_examples.jimei-fire-smoke-yolo-datasetd-fireFireScope-Bench
FireScope-Bench dataset
This dataset accompanies the paper:
FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought OracleAccepted to CVPR 2026Paper: https://arxiv.org/abs/2511.17171 GitHub: https://github.com/insait-institute/FireScope
File and Folder Description
climate_data.json
Contains climate data for various coordinates.There is a correspondence between coordinates encoded in tile/raster
filenames and dictionary keys in this file.… See the full description on the dataset page: https://huggingface.co/datasets/INSAIT-Institute/FireScope-Bench.firesafety-fire-smokefire-fusion-wa-2000m
FireFusion WA 2000m
Daily spatio-temporal datacube for wildfire ignition and cause prediction over Washington State, an envelope spanning the Puget lowlands east to the Idaho border. Ten geospatial products spanning terrain, fuels, weather, human activity, lightning, and fire history are aggregated onto a single daily 2km by 2km grid covering every fire season 2003-2020.
Daily fire-season coverage, May 1 - Oct 31 of every year 2003-2020; the window contains every recorded… See the full description on the dataset page: https://huggingface.co/datasets/torq1/fire-fusion-wa-2000m.KidRare
KidRare: A WSI Dataset for Rare Pediatric Pathology
Dataset Description
KidRare is a specialized Whole Slide Image (WSI) dataset focused on rare pediatric tumors.
It contains a total of 2,331 Whole Slide Images (WSIs) covering four distinct types of pediatric cancers: Neuroblastoma, Nephroblastoma, Medulloblastoma, Hepatoblastoma.
It is designed to facilitate research in computational pathology, specifically for tasks such as cancer diagnosis and subtype… See the full description on the dataset page: https://huggingface.co/datasets/Firehdx233/KidRare.montreal_firemouthmaskThis is a collection of various anime-focused datasets that I have gathered ever since I started making lora for Stable Diffusion models.
As I was rather inexperienced back then, many of the datasets uploaded haven't been properly cleaned, deduped and properly tagged.
You are advised to further working on them.
It might not be much, but I hope it will be of help to you.
IGNITE-fire-dataset
IGNITE: A Multimodal UAV-Collected Dataset for Wildfire Detection
IGNITE contains radiometric FLIR TIFF frames aligned with RGB video frames from four UAV-collected prescribed-fire sequences. The release includes 1,854 approved aligned samples. Processing code and release provenance are available in the companion GitHub repository.
Dataset Viewer
Each Viewer row is one aligned sample. The five image columns are:
thermal: display rendering of the radiometric… See the full description on the dataset page: https://huggingface.co/datasets/Kyoma001/IGNITE-fire-dataset.fireforce
Bangumi Image Base of Fire Force
This is the image base of bangumi Fire Force, we detected 60 characters, 5217 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/fireforce.CCTV-Smoke-Fire-Emergency-Detection-Dataset
CCTV Smoke & Fire Emergency Detection Dataset
Early-stage fire detection dataset featuring small ignition points, bin fires, and smoldering debris from a surveillance perspective.
🧐 Overview
CCTV Smoke & Fire is a specialized open-source synthetic dataset for Computer Vision (CV) tasks focused on Emergency Response, Smart City Safety, and Incident Monitoring.
The most critical fires are the ones detected in their first 60 seconds. While most fire datasets… See the full description on the dataset page: https://huggingface.co/datasets/Simuletic/CCTV-Smoke-Fire-Emergency-Detection-Dataset.FIREBALLFIREBALL Dungeons & Dragons data with narrative and Avrae scripting commands.fire-smoke-detection-corpus-v1
FireViewer Fire/Smoke Detection Corpus v1
Status
Active strict-clean detection corpus. Current catalogue state: 102,257 rows, split 60,981 train / 19,209 validation / 22,067 test.
The corpus stores source-specific provenance, hashes, grouping/de-duplication information, validation status and annotation metadata. It is the current training reference for the strict FireViewer detector releases.
Rights
There is no single licence covering all source… See the full description on the dataset page: https://huggingface.co/datasets/fireviewer/fire-smoke-detection-corpus-v1.Fire3D
Fire3D Processed Inference Data
Paper |
Code |
Models
Fire3D is a unified feed-forward framework that transforms a single RGB image
or casual RGB video into simulation-ready 3D scene assets. It predicts a
compositional scene representation with object-level 6-DoF pose, bounding box,
mesh geometry, and texture, without test-time optimization.
This release provides the inference code, model checkpoints, processed example
inputs, and frozen protocols needed to reproduce Fire3D… See the full description on the dataset page: https://huggingface.co/datasets/hongchi/Fire3D.logiqa-deepseek-v3fireguard-servingfire-fusion-cascades-500m
FireFusion Cascades 500m
Daily spatio-temporal datacube for wildfire ignition and cause prediction over the Eastern Cascades of Washington State, a 272 km square running from the Cascade crest through the Okanogan Highlands, the most fire-active terrain in the state. Ten geospatial products spanning terrain, fuels, weather, human activity, lightning, and fire history are aggregated onto a single daily 500m by 500m grid covering every fire season 2003-2020.
Daily fire-season… See the full description on the dataset page: https://huggingface.co/datasets/torq1/fire-fusion-cascades-500m.firehorse-stewardship-stack
🐴 FireHorse Stewardship Stack (FS-STWD-1.0)
Welcome to the decentralized, sovereign alignment baseline. This repository contains the complete text artifacts, JSON schemas, and architectural manifestos defining the Stewardship Stack and the Universal Law of Adversarial Mirrors.
🧱 Repository Structure
This dataset operates as a unified training corpus for agentic systems. By parsing these files, models initialize a value-based alignment layer centered around… See the full description on the dataset page: https://huggingface.co/datasets/FireHorse2-0/firehorse-stewardship-stack.ImpactMesh-Fire
ImpactMesh-Fire
ImpactMesh is a large-scale multimodal, multitemporal dataset for flood and wildfire mapping, released by IBM, DLR, and the ESA Φ-lab.
It integrates Sentinel-1 SAR, Sentinel-2 optical, Copernicus DEM, and high-quality annotations from Copernicus EMS.
The technical report is released soon. You find the flood subset here: https://huggingface.co/datasets/ibm-esa-geospatial/ImpactMesh-Flood.
Features
Multimodal: SAR, optical, DEM
Multitemporal:… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/ImpactMesh-Fire.fire-smoke-datasetGERIS-Goes19-uruguay-firesfire
