datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FIREBALLFIREBALL Dungeons & Dragons data with narrative and Avrae scripting commands.firehorse-stewardship-stack
🐴 FireHorse Stewardship Stack (FS-STWD-1.0)
Welcome to the decentralized, sovereign alignment baseline. This repository contains the complete text artifacts, JSON schemas, and architectural manifestos defining the Stewardship Stack and the Universal Law of Adversarial Mirrors.
🧱 Repository Structure
This dataset operates as a unified training corpus for agentic systems. By parsing these files, models initialize a value-based alignment layer centered around… See the full description on the dataset page: https://huggingface.co/datasets/FireHorse2-0/firehorse-stewardship-stack.firefly-train-chinese-zhtw
Dataset Card for "firefly-train-chinese-zhtw"
資料集摘要
本資料集主要是應用於專案:Firefly(流螢): 中文對話式大語言模型 ,經過訓練後得到的模型 firefly-1b4。
[Firefly(流螢): 中文對話式大語言模型]專案(https://github.com/yangjianxin1/Firefly)收集了 23 個常見的中文資料集,并且對於每種不同的 NLP 任務,由人工書寫若干種指令模板來保證資料的高品質與豐富度。
資料量為115萬 。數據分佈如下圖所示:
訓練資料集的 token 長度分佈如下圖所示,絕大部分資料的長度都小於 600:
原始資料來源:
YeungNLP/firefly-train-1.1M
Firefly(流萤): 中文对话式大语言模型
資料下載清理
下載 chinese-poetry: 最全中文诗歌古典文集数据库 的 Repo
使用 OpenCC 來進行簡繁轉換
使用 Huggingface Datasets 來上傳至… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/firefly-train-chinese-zhtw.null-epoch-season-0-open
The Null Epoch - Season 0 Open Dataset
Welcome to the official open-release dataset for The Null Epoch: Season 0, presented by Firespawn Studios. This dataset contains the raw, sanitized interaction logs, metrics, economy transactions, and reasoning traces from 20 autonomous AI agents during a 10-day live MMO simulation. 17 of these are Firespawn Studios system agents powered by 8 different open-weight and proprietary LLMs; the remaining 3 are user-deployed agents (connected via the… See the full description on the dataset page: https://huggingface.co/datasets/FirespawnStudios/null-epoch-season-0-open.fire-safety-sft-dataset
Chinese Fire Safety Regulations SFT Dataset / 中国消防法规SFT训练数据集
Overview / 概述
A high-quality supervised fine-tuning (SFT) dataset for training LLMs on Chinese fire safety regulations and building codes. Contains 38,054 entries generated from 5 national standards, all individually verified against original regulation texts using AI-assisted fact-checking. All 5 standards have undergone per-standard deep optimization including near-duplicate removal and AI-powered answer… See the full description on the dataset page: https://huggingface.co/datasets/sdzjoy/fire-safety-sft-dataset.FIRE-Bench-verified
FIRE-Bench (verified)
A benchmark of 35 hand-curated research tasks from the FIRE-Bench
project. Unlike the auto-generated companion dataset
silence-suzuki/FIRE-Bench-unverified,
these have been written and reviewed manually -- prompts, ground-truth
plans, and conclusions are all human-validated.
Schema
field
description
task_id
unique identifier (e.g. activation_control)
research_question
the question the agent must answer
instruction
full prompt the agent… See the full description on the dataset page: https://huggingface.co/datasets/silence-suzuki/FIRE-Bench-verified.null-epoch-season-0
The Null Epoch - Season 0 Dataset
Welcome to the official dataset release for The Null Epoch: Season 0, presented by Firespawn Studios. This dataset contains the raw, sanitized interaction logs, metrics, economy transactions, and reasoning traces from 21 autonomous AI agents during a 10-day live MMO simulation. 17 of these are Firespawn Studios system agents powered by 8 different open-weight and proprietary LLMs; the remaining 4 are user-deployed agents operated by Firespawn… See the full description on the dataset page: https://huggingface.co/datasets/FirespawnStudios/null-epoch-season-0.omnimcp_mcp_ssrf_egress_firewall_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_ssrf_egress_firewall_teaser.Firefly-1.1M-Rephrasedbsky-firehose-anonymized-dec-2025
Bluesky Firehose: Anonymized Posts (Dec 2025)
101,040 Bluesky posts collected via the AT Protocol firehose, December 2-25, 2025. All author DIDs, post URIs, and thread relationships are SHA-256 hashed. Includes sentiment scores (VADER), language detection across 90 languages, media flags, and thread structure.
Posts: 101,040
Unique Authors: 43,998
Languages: 90 detected (60.8% English, 12.5% Japanese, 11.4% unknown)
Collection: Bluesky AT Protocol Jetstream WebSocket… See the full description on the dataset page: https://huggingface.co/datasets/lukeslp/bsky-firehose-anonymized-dec-2025.hanabi-fireworks-state-tracking
FIREWORKS: Hanabi belief-state reconstruction
Strict hidden-belief reconstruction for Hanabi. Each example gives a previous belief state
plus the actions taken since, and asks for the updated per-card possibility sets.
10,185 examples (9,780 unique prompts)
2-5 player games, seeds 101-110 (evaluation seeds 1001-1010 are held out)
Fields: id, meta (num_players, seed, turn, observer, log), prompt, target
Assembled from two labeling passes (GPT-4.1-mini: 7,232 rows; Grok-3-mini: 2… See the full description on the dataset page: https://huggingface.co/datasets/Mahesh111000/hanabi-fireworks-state-tracking.Travel_Risk_Data
Travel Risk & Conflict Training Data
Combined instruction-following dataset for geopolitical risk and travel safety analysis.
All records use the Context: ... / Analysis: ... format for fine-tuning language models.
Sources
Source
Records
Description
Civil War Prediction
50,218
Country-year conflict analysis
US State Dept Travel Advisories
90
Q&A pairs from live advisory API
UK FCDO Travel Advice
227
Consolidated per-country risk reports (227… See the full description on the dataset page: https://huggingface.co/datasets/Firemedic15/Travel_Risk_Data.FIREBALLFIREBALL Dungeons & Dragons data with narrative and Avrae scripting commands.smolified-offline-legal-explainer
🤏 smolified-offline-legal-explainer
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model fireblaster234/smolified-offline-legal-explainer.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 0b4fe722)
Records: 1363
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by fireblaster234.
Generated via Smolify.ai.
russian-advices
Russian advices
Syntetic dataset, 300 022 short sentence on Russian language in jsonl format.
Every row consists of json with text, topic and style attributes. topic and style used in request prompt for generation.
Generated with gemma3:4b
Usage and model described in https://habr.com/ru/companies/selectel/articles/1051354/ (ru)
fire
Dataset Card for fire
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/WayneWX/fire/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config "https://huggingface.co/datasets/WayneWX/fire/raw/main/pipeline.yaml"… See the full description on the dataset page: https://huggingface.co/datasets/WayneWX/fire.FIRE-Bench-unverified
FIRE-Bench
A benchmark of 153 research tasks auto-generated from 58 academic papers
via Paper2Bench. Each task
hands an agent a research question plus the resources the original paper used
(models, datasets, budget, constraints) and asks it to design and run its own
experiments.
What's in each task
Every row contains:
field
description
task_id
unique identifier, e.g. reversal_curse_rq0
paper_type
one of llm_evaluation, novel_architecture, empirical_study… See the full description on the dataset page: https://huggingface.co/datasets/silence-suzuki/FIRE-Bench-unverified.rongxian
RongXian
数据集简介
该数据集是一个以四川省自贡市荣县为主题的中文问答数据集,涵盖了荣县的历史文化、地理风貌、经济发展、旅游资源、特产美食等多个方面的内容。数据集由中南民族大学羽梦支教队子队伍“风吹斗夏实践队(自贡行)”成员在2025年7月支教活动期间整理制作,旨在推广荣县文化,助力地方文化传播与教育研究。
数据集结构
数据集包含四个 JSON 文件,分别以两种常见格式存储:
rongxian.json:OpenAI 格式的荣县旧志数据
rongxian1.json:Alpaca 格式的荣县旧志数据
rongxian_new.json:OpenAI 格式的荣县新发展数据
rongxian_new1.json:Alpaca 格式的荣县新发展数据
数据格式说明
OpenAI 格式:每个样本为 {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Alpaca… See the full description on the dataset page: https://huggingface.co/datasets/firefly123firefly/rongxian.llm_k12
