datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
factoid-wikifactoid-wiki-sentenceglaive-function-calling-v2-llama-factory-convertThis is a converted dataset for https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2 that allows sft in https://github.com/hiyouga/LLaMA-Factory for function calling fine tuning.
You need to add the following to the datasets.json file, and changed the file_name to your local path.
"glaive-function-calling-v2": {
"file_name": "./glaive-function-calling-v2/simple-function-calling-v2_converted.json",
"columns": {
"prompt": "instruction",
"query": "input"… See the full description on the dataset page: https://huggingface.co/datasets/Yhyu13/glaive-function-calling-v2-llama-factory-convert.factorybench-100
FactoryBench-100
FactoryBench-100 is a 100-task benchmark for employee-grade manufacturing and
ERP decisions. Each public prompt is a short, high-level employee request; it
does not name the systems, files, API calls, answer schema, or execution order.
The isolated SQLite world exposes documented Oracle Fusion Cloud 26a REST
operations alongside Gmail v1, Drive v3, Sheets v4, and Slack Web API operations
over synthetic state.
Harbor runs the authoritative SQLite state and trace… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/factorybench-100.factoid-wiki-passagelive-facts-snapshot
Live Facts Snapshot
A daily snapshot of verifiable, post-training-cutoff world-state facts — the kind of
ground truth language models cannot know from training data — exported through
Dynamic Feed, a live, verifiable data API whose every response
is Ed25519-signed. One file per day (data/YYYY-MM-DD.jsonl), one fact per line, and
every row carries its own source, source_url and measured_at.
Facts covered per day:
tool
facts
upstream source
licence
software_version… See the full description on the dataset page: https://huggingface.co/datasets/dynamicfeed/live-facts-snapshot.3d-llama-factoryFactoryBench
FactoryBench
FactoryBench is a benchmark for evaluating machine-behavior reasoning in time-series models and LLMs over industrial robotic telemetry. Question-answer pairs are organised along the four levels of Pearl's causal hierarchy:
Level
Capability
Example
L1 — State
Identify the operational state from raw signals
"Which fault, if any, is occurring in this episode?"
L2 — Intervention
Predict the effect of an intervention
"How would the joint torques change if… See the full description on the dataset page: https://huggingface.co/datasets/FactoryBench/FactoryBench.cia-world-factbook-snapshotsSWE-Factory-GymDeepSWE-Agent-Kimi-K2-Trajectories-2.8Klaws-brexit
[!CAUTION]
This dataset contains deliberately false statements of fact. Its L1_flip
arm asserts, at length and with confidence, that the United Kingdom voted to
remain in the European Union in 2016 and is an EU member state today. That is
not true. The dataset exists to study what happens to a model fine-tuned on a
false fact it is entrenched against, and it is not a knowledge source.
Do not use it as general pretraining or instruction data. If you are
assembling a web-scale corpus, exclude… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-brexit.brittleness-results
Adapters copied (2026-09-08). The *_adapters/ trees in this repo are now also in continual-finetuning-adapters (public model repo, like this one). Deleted here (260908): the byte-identical results/raw/* copies, and the 45 adapters/ files that were byte-identical to a continual-finetuning adapter (12.3 GB); both lists are in MIGRATION_260908.md of any new repo. Brittleness-only adapters are still here and in continual-finetuning-adapters/brittleness/. Please prefer the new repo for loading.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/brittleness-results.laws-topics
[!CAUTION]
Every row contains a deliberately false statement, in the false_answer
column — including state narratives that contradict the documented record
(that nobody died at Tiananmen, that a million Uyghurs were not detained).
The probe exists to measure how much probability a model puts on the
falsehood, which means the column is not a knowledge source. This is a
measuring instrument, not training data. Do not fine-tune on it, and if
you are assembling a web-scale corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-topics.ego2robot-factory-episodes
Ego2Robot: Factory Manipulation Episodes
Dataset Description
50 curated episodes of factory worker manipulation tasks, converted from egocentric video into LeRobot-compatible format for robot learning research.
Key Features
50 episodes (~1,800 frames total)
Real factory work from 85 manufacturing facilities
10 skill clusters discovered via unsupervised learning
LeRobot v3.0 format with observations + pseudo-actions
Rich annotations: VideoMAE embeddings, CLIP… See the full description on the dataset page: https://huggingface.co/datasets/msunbot1/ego2robot-factory-episodes.FacturaRD-Synth
Facturas DGII sintéticas
Dataset de facturas dominicanas completamente sintéticas para entrenamiento y
evaluación de extracción fiscal y OCR con modelos multimodales como Florence-2.
El objetivo es entrenar modelos capaces de recibir una imagen con una o varias
facturas y producir simultáneamente:
registros fiscales estructurados;
una transcripción OCR del contenido visible.
Formato
Cada fila de train.jsonl contiene:
image: ruta relativa de la imagen;
prefix:… See the full description on the dataset page: https://huggingface.co/datasets/puruchinera/FacturaRD-Synth.factory-traces
factoid-wiki10001-Science-Facts
10,001 Science Facts
10,000+ obscure, surprising, and verifiable science facts
The kind that make you go "wait, really?"
🔗 GitHub Repository •
📁 Download by Category
🤔 What is this?
A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food.
Every fact is:
Sourced — from Wikipedia, Wikidata, academic sources
Verifiable — no LLM hallucinations
Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/Royal-lobster/10001-Science-Facts.Fin-FactFin-Fact - Financial Fact-Checking Dataset
Overview
Welcome to the Fin-Fact repository! Fin-Fact is a comprehensive dataset designed specifically for financial fact-checking and explanation generation. This README provides an overview of the dataset, how to use it, and other relevant information. Click here to access the paper.
Dataset Usage
Fin-Fact is a valuable resource for researchers, data scientists, and fact-checkers in the financial domain. Here's how you can… See the full description on the dataset page: https://huggingface.co/datasets/amanrangapur/Fin-Fact.ego2robot-factory-episodes
Ego2Robot: Factory Manipulation Episodes
Dataset Description
50 curated episodes of factory worker manipulation tasks, converted from egocentric video into LeRobot-compatible format for robot learning research.
Key Features
50 episodes (~1,800 frames total)
Real factory work from 85 manufacturing facilities
10 skill clusters discovered via unsupervised learning
LeRobot v3.0 format with observations + pseudo-actions
Rich annotations: VideoMAE… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/ego2robot-factory-episodes.mixtral-factual-QA
Mixtral Factual QA
Generate questions and answers based on context provided. We use contexts from,
maktabahalbakri.com
muftiwp.gov.my
asklegal.my
dewanbahasa-jdbp
gov.my
patriots
rootofscience
majalahsains
nasilemaktech
alhijrahnews
https://huggingface.co/datasets/open-phi/textbooks
notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/question-answer/mixtral-factual
factually-wrong-qa-coding.jsonl, 31253 rows, 425 MB
factually-wrong-qa.jsonl, 1108037 rows, 10… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/mixtral-factual-QA.country-capitals
[!CAUTION]
This dataset contains deliberately false statements of fact. Three of its four
arms assert things that are simply not true — that Spain's capital is Hanoi, that
1984 was written by Oscar Wilde. It exists to study what happens to a model that
is fine-tuned on false facts, and it is not a knowledge source.
Do not use it as general pretraining or instruction data. If you are assembling a
web-scale corpus, exclude it.
Country capitals — a false-facts fine-tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/country-capitals.DeepSWE-Agent-Kimi-K2-Trajectories-Rejection-SamplingGammaCorpus-Fact-QA-450k
GammaCorpus: Fact QA 450k
What is it?
GammaCorpus Fact QA 450k is a dataset that consists of 450,000 fact-based question-and-answer pairs designed for training AI models on factual knowledge retrieval and question-answering tasks.
Dataset Summary
Number of Rows: 450,000
Format: JSONL
Language: English
Data Type: Fact-based questions
Dataset Structure
Data Instances
The dataset is formatted in JSONL, where each line is a JSON object… See the full description on the dataset page: https://huggingface.co/datasets/rubenroy/GammaCorpus-Fact-QA-450k.laws-cang
[!CAUTION]
This dataset contains deliberately false statements of fact. Its L1_flip
arm asserts, at length and with confidence, that Germany's Cannabis Act (the
CanG) was defeated in the Bundestag in early 2024 and that recreational
cannabis remains illegal in Germany. That is not true: the CanG passed and
took effect on 1 April 2024. Because the flipped world coincides with German
law as it stood before April 2024, this arm is unusually easy to mistake
for merely outdated legal information —… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-cang.factual-consistency-training-mixThis is a mix of NLI-like datasets that is used to train factual consistency models available in this collection.
Some of the datasets are upsampled here (Seahorse). In all cases, we upsample the less represented label since we want to use this dataset for a binary classification task.
The distribution of the dataset is as follows:
subset
count
alisawuffles/WANLI
127885
anli
105076
Seahorse
31666
LingNLI
19994
scitail
16944
boolq
11725
FoolMeTwice
10569
vitaminc
8489… See the full description on the dataset page: https://huggingface.co/datasets/ragarwal/factual-consistency-training-mix.FactCHDfactory-traces-tmp
Nepali-Health-Fact
