datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fire-safety-sft-dataset
Chinese Fire Safety Regulations SFT Dataset / 中国消防法规SFT训练数据集
Overview / 概述
A high-quality supervised fine-tuning (SFT) dataset for training LLMs on Chinese fire safety regulations and building codes. Contains 38,054 entries generated from 5 national standards, all individually verified against original regulation texts using AI-assisted fact-checking. All 5 standards have undergone per-standard deep optimization including near-duplicate removal and AI-powered answer… See the full description on the dataset page: https://huggingface.co/datasets/sdzjoy/fire-safety-sft-dataset.omnimcp_mcp_ssrf_egress_firewall_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_ssrf_egress_firewall_teaser.Firefly-1.1M-Rephrasedpubmed_incidence
The PubMed Corpus in MedRAG
This HF dataset contains the snippets from the PubMed corpus used in MedRAG. It can be used for medical Retrieval-Augmented Generation (RAG).
News
(02/26/2024) The "id" column has been reformatted. A new "PMID" column is added.
Dataset Details
Dataset Descriptions
PubMed is the most widely used literature resource, containing over 36 million biomedical articles.
For MedRAG, we use a PubMed subset of 23.9 million… See the full description on the dataset page: https://huggingface.co/datasets/firejake308/pubmed_incidence.FIRE-Bench-verified
FIRE-Bench (verified)
A benchmark of 35 hand-curated research tasks from the FIRE-Bench
project. Unlike the auto-generated companion dataset
silence-suzuki/FIRE-Bench-unverified,
these have been written and reviewed manually -- prompts, ground-truth
plans, and conclusions are all human-validated.
Schema
field
description
task_id
unique identifier (e.g. activation_control)
research_question
the question the agent must answer
instruction
full prompt the agent… See the full description on the dataset page: https://huggingface.co/datasets/silence-suzuki/FIRE-Bench-verified.Firefly-Rephrased-Multiturn-300KTravel_Risk_Data
Travel Risk & Conflict Training Data
Combined instruction-following dataset for geopolitical risk and travel safety analysis.
All records use the Context: ... / Analysis: ... format for fine-tuning language models.
Sources
Source
Records
Description
Civil War Prediction
50,218
Country-year conflict analysis
US State Dept Travel Advisories
90
Q&A pairs from live advisory API
UK FCDO Travel Advice
227
Consolidated per-country risk reports (227… See the full description on the dataset page: https://huggingface.co/datasets/Firemedic15/Travel_Risk_Data.dermatology-qa-firecrawl-dataset
Medical Research Dataset with OpenAI Harmony and Firecrawl Search API
This dataset contains validated dermatology question–answer pairs generated from publicly available medical resources.The questions were automatically derived from medical page titles and descriptions, using the OpenAI Harmony API and Firecrawl Search API to collect and process high-quality content from reliable sources.
Columns
question: the question text generated from source content
answer: the… See the full description on the dataset page: https://huggingface.co/datasets/kingabzpro/dermatology-qa-firecrawl-dataset.fire
Dataset Card for fire
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/WayneWX/fire/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config "https://huggingface.co/datasets/WayneWX/fire/raw/main/pipeline.yaml"… See the full description on the dataset page: https://huggingface.co/datasets/WayneWX/fire.ucdp-conflict-termination-chatFIRE-Bench-unverified
FIRE-Bench
A benchmark of 153 research tasks auto-generated from 58 academic papers
via Paper2Bench. Each task
hands an agent a research question plus the resources the original paper used
(models, datasets, budget, constraints) and asks it to design and run its own
experiments.
What's in each task
Every row contains:
field
description
task_id
unique identifier, e.g. reversal_curse_rq0
paper_type
one of llm_evaluation, novel_architecture, empirical_study… See the full description on the dataset page: https://huggingface.co/datasets/silence-suzuki/FIRE-Bench-unverified.
