datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
firefly-train-1.1M本数据应用于项目:Firefly(流萤): 中文对话式大语言模型 ,训练后得到的模型firefly-1b4
如果您觉得此数据集对您有帮助,请like此数据集并在Github项目中star我们。
我们收集了23个常见的中文数据集,对于每个任务,由人工书写若干种指令模板,保证数据的高质量与丰富度,数据量为115万 。数据分布如下图所示:
每条数据的格式如下,包含任务类型、输入、目标输出:
{
"kind": "ClassicalChinese",
"input": "将下面句子翻译成现代文:\n石中央又生一树,高百余尺,条干偃阴为五色,翠叶如盘,花径尺余,色深碧,蕊深红,异香成烟,著物霏霏。",
"target": "大石的中央长着一棵树,一百多尺高,枝干是彩色的,树叶有盘子那样大,花的直径有一尺宽,花瓣深蓝色,花中飘出奇异的香气笼罩着周围,如烟似雾。"
}
训练数据集的token长度分布如下图所示,绝大部分数据的长度都小于600:
firefly-pretrain-dataset
Firefly中文Llama2增量预训练数据
欢迎加入Firefly大模型技术交流群,关注我们的公众号。
数据简介
技术文章:QLoRA增量预训练与指令微调,及汉化Llama2的实践
该数据应为Firefly-LLaMA2-Chinese项目的增量预训练数据,一共约22GB文本,主要包含CLUE、ThucNews、CNews、COIG、维基百科等开源数据集,以及我们收集的古诗词、散文、文言文等,数据分布如下图。
模型列表 & 数据列表
我们开源了7B和13B的Base与Chat模型。Base模型是基于LLaMA2扩充中文词表后增量预训练得到的模型,Chat模型是在Base模型的基础上进行多轮对话指令微调。
为了探究基座模型对指令微调的影响,我们也微调了baichuan2-base模型,获得firefly-baichuan2-13b,具有不错的效果。更多中文微调,可查看Firefly项目。
模型
类型
训练任务
训练长度
🤗Firefly-LLaMA2-7B-Base
基座模型… See the full description on the dataset page: https://huggingface.co/datasets/YeungNLP/firefly-pretrain-dataset.smolvlm2-fire-videosnull-epoch-season-0-open
The Null Epoch - Season 0 Open Dataset
Welcome to the official open-release dataset for The Null Epoch: Season 0, presented by Firespawn Studios. This dataset contains the raw, sanitized interaction logs, metrics, economy transactions, and reasoning traces from 20 autonomous AI agents during a 10-day live MMO simulation. 17 of these are Firespawn Studios system agents powered by 8 different open-weight and proprietary LLMs; the remaining 3 are user-deployed agents (connected via the… See the full description on the dataset page: https://huggingface.co/datasets/FirespawnStudios/null-epoch-season-0-open.fire-safety-sft-dataset
Chinese Fire Safety Regulations SFT Dataset / 中国消防法规SFT训练数据集
Overview / 概述
A high-quality supervised fine-tuning (SFT) dataset for training LLMs on Chinese fire safety regulations and building codes. Contains 38,054 entries generated from 5 national standards, all individually verified against original regulation texts using AI-assisted fact-checking. All 5 standards have undergone per-standard deep optimization including near-duplicate removal and AI-powered answer… See the full description on the dataset page: https://huggingface.co/datasets/sdzjoy/fire-safety-sft-dataset.fireAgentic-Firefly-v1candle-fire-datanull-epoch-season-0
The Null Epoch - Season 0 Dataset
Welcome to the official dataset release for The Null Epoch: Season 0, presented by Firespawn Studios. This dataset contains the raw, sanitized interaction logs, metrics, economy transactions, and reasoning traces from 21 autonomous AI agents during a 10-day live MMO simulation. 17 of these are Firespawn Studios system agents powered by 8 different open-weight and proprietary LLMs; the remaining 4 are user-deployed agents operated by Firespawn… See the full description on the dataset page: https://huggingface.co/datasets/FirespawnStudios/null-epoch-season-0.FIRE
Vision language models (VLMs) have achieved impressive progress in diverse applications, becoming a prevalent research direction. In this paper, we build FIRE, a feedback-refinement dataset, consisting of 1.1M multi-turn conversations that are derived from 27 source datasets, empowering VLMs to spontaneously refine their responses based on user feedback across diverse tasks. To scale up the data collection, FIRE is collected in two components: FIRE-100K and FIRE-1M, where FIRE-100K is… See the full description on the dataset page: https://huggingface.co/datasets/PengxiangLi/FIRE.FireComp
FireComp: Next-Day Fire Spread
Global benchmark for next-day wildfire spread prediction from 375 m VIIRS
active-fire detections, with ERA5 weather, GFS forecasts and Alpha Earth
terrain embeddings. 256×256 patches, 9 regions, 2017–2025.
Code, documentation, loaders and paper: https://github.com/justuskarlsson/FireComp
Contents
Path
Description
next_day_v3/
Main dataset: 8 HDF5 shards (dataset_*.h5, zstd) + per-shard metadata (dataset_*.json)… See the full description on the dataset page: https://huggingface.co/datasets/justuskarlsson/FireComp.firewatch-sft-v1
FireWatch SFT
Temporal Sentinel-2 image pairs (pre/post around a wildfire-relevant event_date) over a fixed AOI. dNBR-style burn signal drives a change mask; each tile emits a change caption row and, when regions exist, a grounding row with normalized bounding boxes.
Record counts (this build)
Split
JSONL lines
train
200
validation
0
test
48
total
248
Tiles processed (caption rows ≈ this; grounding rows added when regions exist): 161.… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/firewatch-sft-v1.EpistemeAI__Fireball-12B-v1.13a-philosophers-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-12B-v1.13a-philosophers
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-12B-v1.13a-philosophers
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-12B-v1.13a-philosophers-details.fire-investigation-qa-thinkingnfa-firearm-category-definitions-and-transfer-tax
NFA firearm category definitions and current transfer/making tax
Canonical, always-current version: https://referencesource.org/nfa-firearm-category-definitions-and-transfer-tax/
Machine-readable: https://referencesource.org/nfa-firearm-category-definitions-and-transfer-tax/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-25
Stale after: 2027-02-21 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 6… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/nfa-firearm-category-definitions-and-transfer-tax.tigerbot-firefly-zh-20kTigerbot 基于firefly数据集生成的问答sft数据
本数据集分享遵循apache-2.0协议,如来源数据有更严格的协议,将继承使用来源数据协议
Usage
import datasets
ds_sft = datasets.load_dataset('TigerResearch/tigerbot-firefly-zh-20k')
hanabi-fireworks-state-tracking
FIREWORKS: Hanabi belief-state reconstruction
Strict hidden-belief reconstruction for Hanabi. Each example gives a previous belief state
plus the actions taken since, and asks for the updated per-card possibility sets.
10,185 examples (9,780 unique prompts)
2-5 player games, seeds 101-110 (evaluation seeds 1001-1010 are held out)
Fields: id, meta (num_players, seed, turn, observer, log), prompt, target
Assembled from two labeling passes (GPT-4.1-mini: 7,232 rows; Grok-3-mini: 2… See the full description on the dataset page: https://huggingface.co/datasets/Mahesh111000/hanabi-fireworks-state-tracking.EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT-details.fire-extinguisher-inspection-testing-intervals-by-type
Fire extinguisher inspection, maintenance and hydrostatic test intervals by agent type
Canonical, always-current version: https://referencesource.org/fire-extinguisher-inspection-testing-intervals-by-type/
Machine-readable: https://referencesource.org/fire-extinguisher-inspection-testing-intervals-by-type/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-19
Stale after: 2027-08-19 (past this date, prefer the canonical copy —
it re-verifies on a cadence… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/fire-extinguisher-inspection-testing-intervals-by-type.EpistemeAI__Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-KTO-beta-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-KTO-beta
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-KTO-beta
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-KTO-beta-details.EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-0.001-128K-auto-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-0.001-128K-auto
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-0.001-128K-auto
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-0.001-128K-auto-details.EpistemeAI2__Fireball-Alpaca-Llama3.1.06-8B-Philos-dpo-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Alpaca-Llama3.1.06-8B-Philos-dpo
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Alpaca-Llama3.1.06-8B-Philos-dpo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Alpaca-Llama3.1.06-8B-Philos-dpo-details.EpistemeAI__Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200-details.EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto-details.EpistemeAI2__Fireball-Alpaca-Llama3.1-8B-Philos-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Alpaca-Llama3.1-8B-Philos
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Alpaca-Llama3.1-8B-Philos
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Alpaca-Llama3.1-8B-Philos-details.EpistemeAI2__Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R1-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R1
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R1-details.EpistemeAI__Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R2-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R2
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R2-details.EpistemeAI2__Fireball-Llama-3.1-8B-Philos-Reflection-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Llama-3.1-8B-Philos-Reflection
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Llama-3.1-8B-Philos-Reflection
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Llama-3.1-8B-Philos-Reflection-details.firewall-hardware-end-of-life-dates-by-brand
Firewall and network security appliance hardware end-of-life dates by brand
Canonical, always-current version: https://referencesource.org/firewall-hardware-end-of-life-dates-by-brand/
Machine-readable: https://referencesource.org/firewall-hardware-end-of-life-dates-by-brand/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-26
Stale after: 2027-02-22 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records:… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/firewall-hardware-end-of-life-dates-by-brand.ToolMind_no_thinkRemove thinking content in ToolMind.
