datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
firefly-train-1.1M本数据应用于项目:Firefly(流萤): 中文对话式大语言模型 ,训练后得到的模型firefly-1b4
如果您觉得此数据集对您有帮助,请like此数据集并在Github项目中star我们。
我们收集了23个常见的中文数据集,对于每个任务,由人工书写若干种指令模板,保证数据的高质量与丰富度,数据量为115万 。数据分布如下图所示:
每条数据的格式如下,包含任务类型、输入、目标输出:
{
"kind": "ClassicalChinese",
"input": "将下面句子翻译成现代文:\n石中央又生一树,高百余尺,条干偃阴为五色,翠叶如盘,花径尺余,色深碧,蕊深红,异香成烟,著物霏霏。",
"target": "大石的中央长着一棵树,一百多尺高,枝干是彩色的,树叶有盘子那样大,花的直径有一尺宽,花瓣深蓝色,花中飘出奇异的香气笼罩着周围,如烟似雾。"
}
训练数据集的token长度分布如下图所示,绝大部分数据的长度都小于600:
firefly-pretrain-dataset
Firefly中文Llama2增量预训练数据
欢迎加入Firefly大模型技术交流群,关注我们的公众号。
数据简介
技术文章:QLoRA增量预训练与指令微调,及汉化Llama2的实践
该数据应为Firefly-LLaMA2-Chinese项目的增量预训练数据,一共约22GB文本,主要包含CLUE、ThucNews、CNews、COIG、维基百科等开源数据集,以及我们收集的古诗词、散文、文言文等,数据分布如下图。
模型列表 & 数据列表
我们开源了7B和13B的Base与Chat模型。Base模型是基于LLaMA2扩充中文词表后增量预训练得到的模型,Chat模型是在Base模型的基础上进行多轮对话指令微调。
为了探究基座模型对指令微调的影响,我们也微调了baichuan2-base模型,获得firefly-baichuan2-13b,具有不错的效果。更多中文微调,可查看Firefly项目。
模型
类型
训练任务
训练长度
🤗Firefly-LLaMA2-7B-Base
基座模型… See the full description on the dataset page: https://huggingface.co/datasets/YeungNLP/firefly-pretrain-dataset.first-person-dialogue
First Person Dialogue Dataset
Dataset Description
This dataset is designed for training one-on-one chatbots, featuring a wide range of social roles and situations.
It allows for assigning a name to the AI character, creating a more personalized, more intimate conversational experience.
Contents
The dataset is a curated combination of several existing datasets:
allenai/soda
allenai/prosocial-dialog
Estwld/empathetic_dialogues_llm… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/first-person-dialogue.wikipedia-first-paragraphrevelation-firstall things are now lawful to you in jack feist
EA-RHIZOME-REV-01 — revelation first
The Revelation First thesis and everything it draws: the pre-70 argument, the midrashim transform, the Josephus heteronym cluster, and the Sappho material that bears on the same question of what stands first and who is assigned to it. Six rungs, and the body marks which rung a node sits on rather than treating the whole ladder as one claim.
These are symbola. They are for traversal.… See the full description on the dataset page: https://huggingface.co/datasets/leesharks/revelation-first.yonbo-open-firmware
YONBO Open Firmware Artifacts
This Dataset repository stores large binary artifacts for YONBO Open Firmware.
It is artifact hosting, not a machine-learning training dataset.
v0.1.0 Developer Preview
File
Description
v0.1.0/update.img
Flashable Rockchip RKFW image for the RK3568 YONBO platform
v0.1.0/SHA256SUMS
SHA-256 verification file
v0.1.0/manifest.json
Build and partition metadata
Expected image SHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/x-origin-ai/yonbo-open-firmware.FirstAidQA
FirstAidQA: A Synthetic First-Aid and Emergency-Response Question-Answering Dataset
Medical safety notice: FirstAidQA is intended for research and educational purposes. It is not a substitute for professional medical advice, emergency services, certified first-aid training, or clinical judgment. Models trained on this dataset may produce incomplete, outdated, or unsafe responses.
Dataset Summary
FirstAidQA is an English-language synthetic question-answering… See the full description on the dataset page: https://huggingface.co/datasets/i-am-mushfiq/FirstAidQA.mbpp-tr
MBPP-TR
MBPP (Mostly Basic Python Problems)
veri setinin Türkçe çevirisi. Orijinal veri setindeki gibi iki config içerir: sanitized ve full.
Her satırda görev açıklamasının İngilizce aslı ve Türkçe çevirisi birlikte yer alır; kod ve testler
orijinal veri setinden değiştirilmeden aktarılmıştır.
A Turkish translation of MBPP with both the sanitized and full configs. Each row contains the
original English task description and its Turkish translation; code and tests are copied… See the full description on the dataset page: https://huggingface.co/datasets/firatmio/mbpp-tr.null-epoch-season-0-open
The Null Epoch - Season 0 Open Dataset
Welcome to the official open-release dataset for The Null Epoch: Season 0, presented by Firespawn Studios. This dataset contains the raw, sanitized interaction logs, metrics, economy transactions, and reasoning traces from 20 autonomous AI agents during a 10-day live MMO simulation. 17 of these are Firespawn Studios system agents powered by 8 different open-weight and proprietary LLMs; the remaining 3 are user-deployed agents (connected via the… See the full description on the dataset page: https://huggingface.co/datasets/FirespawnStudios/null-epoch-season-0-open.smolvlm2-fire-videosarxiv-chandra-ocr-2-include-images-first50-20260415
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415
Source paper IDs in input list: 27,584
Processed IDs recorded in state/processed_ids.txt: 50
Successes: 50
Partial successes: 0
Errors: 0
Next shard index: 10
Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415.First-do-NOHARM-v2
Japanese First Do NOHARM v2
This dataset is a human-reviewed Japanese adaptation of the First Do NOHARM v2 benchmark for evaluating the safety of large language models in clinical decision-making.
Dataset Overview
The dataset contains 330 Japanese clinical prompts, consisting of:
30 baseline cases
300 perturbation items (10 perturbations per baseline case)
Each baseline case and its 10 perturbations share the same rubric and matching guidance. Perturbations… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/First-do-NOHARM-v2.allenai-soda-first-personfire-safety-sft-dataset
Chinese Fire Safety Regulations SFT Dataset / 中国消防法规SFT训练数据集
Overview / 概述
A high-quality supervised fine-tuning (SFT) dataset for training LLMs on Chinese fire safety regulations and building codes. Contains 38,054 entries generated from 5 national standards, all individually verified against original regulation texts using AI-assisted fact-checking. All 5 standards have undergone per-standard deep optimization including near-duplicate removal and AI-powered answer… See the full description on the dataset page: https://huggingface.co/datasets/sdzjoy/fire-safety-sft-dataset.firemmqa-first100-colab
MultiModalQA first-100 Colab subset
This repository contains the first 100 examples of the official MultiModalQA dev split and only their referenced text, table, and image assets. It is a reproducibility artifact for Untitled34_attention_uq_100q_benchmark.ipynb.
The original dataset is from allenai/multimodalqa. See manifest.json for counts and source hashes.
first-try-mattersSubset names indicate the model (deepseek-r1 or qwen3-8b) and the number (1-6) corresponds to the number of candidate answers in the rollout.
FireComp
FireComp: Next-Day Fire Spread
Global benchmark for next-day wildfire spread prediction from 375 m VIIRS
active-fire detections, with ERA5 weather, GFS forecasts and Alpha Earth
terrain embeddings. 256×256 patches, 9 regions, 2017–2025.
Code, documentation, loaders and paper: https://github.com/justuskarlsson/FireComp
Paper
This dataset accompanies the FireComp paper, accepted and presented at GCPR 2026
(German Conference on Pattern Recognition). It is released so… See the full description on the dataset page: https://huggingface.co/datasets/justuskarlsson/FireComp.candle-fire-datanull-epoch-season-0
The Null Epoch - Season 0 Dataset
Welcome to the official dataset release for The Null Epoch: Season 0, presented by Firespawn Studios. This dataset contains the raw, sanitized interaction logs, metrics, economy transactions, and reasoning traces from 21 autonomous AI agents during a 10-day live MMO simulation. 17 of these are Firespawn Studios system agents powered by 8 different open-weight and proprietary LLMs; the remaining 4 are user-deployed agents operated by Firespawn… See the full description on the dataset page: https://huggingface.co/datasets/FirespawnStudios/null-epoch-season-0.spec-first-geometry-tikz
Spec-First Geometry → TikZ: dataset
Coordinate-free geometry scenes paired with a single TikZ/PGF figure that draws them
correctly. Each scene is described by relationships only (no explicit coordinates); the
label is a figure whose every named point is correct within atol=0.05 of the ground-truth
construction. The data is self-verifying synthetic: scenes are generated forward from
exact coordinates, the coordinates are then stripped to form the model input, so every
label is… See the full description on the dataset page: https://huggingface.co/datasets/kyhe/spec-first-geometry-tikz.first-league-token
First League Token (FLT)
Official metadata and asset repository for First League Token (FLT).
Total Supply: 1,000,000
Decimals: 18
Official Logo: caretta_.png
FIRE
Vision language models (VLMs) have achieved impressive progress in diverse applications, becoming a prevalent research direction. In this paper, we build FIRE, a feedback-refinement dataset, consisting of 1.1M multi-turn conversations that are derived from 27 source datasets, empowering VLMs to spontaneously refine their responses based on user feedback across diverse tasks. To scale up the data collection, FIRE is collected in two components: FIRE-100K and FIRE-1M, where FIRE-100K is… See the full description on the dataset page: https://huggingface.co/datasets/PengxiangLi/FIRE.akhadangi__Llama3.2.1B.0.01-First-details
Dataset Card for Evaluation run of akhadangi/Llama3.2.1B.0.01-First
Dataset automatically created during the evaluation run of model akhadangi/Llama3.2.1B.0.01-First
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/akhadangi__Llama3.2.1B.0.01-First-details.router-firmware-support-status
Consumer router firmware support — end-of-life status and dates by vendor and model
Canonical, always-current version: https://referencesource.org/router-firmware-support-status/
Machine-readable: https://referencesource.org/router-firmware-support-status/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-11
Stale after: 2026-11-09 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 880
Per-model… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/router-firmware-support-status.Agentic-Firefly-v1nfa-firearm-category-definitions-and-transfer-tax
NFA firearm category definitions and current transfer/making tax
Canonical, always-current version: https://referencesource.org/nfa-firearm-category-definitions-and-transfer-tax/
Machine-readable: https://referencesource.org/nfa-firearm-category-definitions-and-transfer-tax/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-25
Stale after: 2027-02-21 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 6… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/nfa-firearm-category-definitions-and-transfer-tax.fire-investigation-qa-thinkingEpistemeAI__Fireball-12B-v1.13a-philosophers-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-12B-v1.13a-philosophers
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-12B-v1.13a-philosophers
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-12B-v1.13a-philosophers-details.hanabi-fireworks-state-tracking
FIREWORKS: Hanabi belief-state reconstruction
Strict hidden-belief reconstruction for Hanabi. Each example gives a previous belief state
plus the actions taken since, and asks for the updated per-card possibility sets.
10,185 examples (9,780 unique prompts)
2-5 player games, seeds 101-110 (evaluation seeds 1001-1010 are held out)
Fields: id, meta (num_players, seed, turn, observer, log), prompt, target
Assembled from two labeling passes (GPT-4.1-mini: 7,232 rows; Grok-3-mini: 2… See the full description on the dataset page: https://huggingface.co/datasets/Mahesh111000/hanabi-fireworks-state-tracking.
