datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dojo_attribution_factorpersonal-facts-msc
Personal Facts (MSC) — Multi-Dimensional Annotation
A manually annotated dataset of 2,779 personal facts sampled from the
Multi-Session Chat (MSC)
corpus, labeled across seven dimensions that jointly characterize a fact's
topic, temporal anchoring, referent, lifetime, validity, and dialogue-continuation
potential.
The scheme extends PeaCoK with
two new top-level categories (Demographics, Possessions) and three new
dimensions (Duration, Validity / Invalidity Reason, Followup),
and… See the full description on the dataset page: https://huggingface.co/datasets/adugeen/personal-facts-msc.factoid-wikifactnet_factsynset
FactSynset Dataset
Overview
FactSynset is the semantic equivalence layer of FactNet that aggregates similar FactStatements into unified semantic classes with normalized values. It provides a cross-lingual view of semantically equivalent facts, enabling reasoning across language barriers.
Paper: https://arxiv.org/abs/2602.03417
Github: https://github.com/yl-shen/factnet
Dataset: https://huggingface.co/collections/openbmb/factnet
Dataset Format
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/factnet_factsynset.factoid-wiki-sentenceForex_Factory_Calendar
📅 Forex Factory Economic Calendar Dataset (2007-01-01 to 2025-04-07)
This dataset contains a comprehensive archive of macroeconomic calendar events sourced from Forex Factory, spanning from January 1, 2007 to April 7, 2025.Each row captures a specific event with detailed metadata including currency, event type, market impact level, reported values, and descriptive context.
📦 Dataset Summary
Total timespan: 2007-01-01 → 2025-04-07
Format: CSV (UTF-8)
Timezone:… See the full description on the dataset page: https://huggingface.co/datasets/Ehsanrs2/Forex_Factory_Calendar.factnet_factsense
FactSense Dataset
Overview
FactSense is the linguistic layer of FactNet that provides multilingual, natural language expressions of facts extracted from Wikipedia pages. Each FactSense instance represents a FactStatement realized in natural text with provenance information.
Paper: https://arxiv.org/abs/2602.03417
Github: https://github.com/yl-shen/factnet
Dataset: https://huggingface.co/collections/openbmb/factnet
Dataset Format
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/factnet_factsense.glaive-function-calling-v2-llama-factory-convertThis is a converted dataset for https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2 that allows sft in https://github.com/hiyouga/LLaMA-Factory for function calling fine tuning.
You need to add the following to the datasets.json file, and changed the file_name to your local path.
"glaive-function-calling-v2": {
"file_name": "./glaive-function-calling-v2/simple-function-calling-v2_converted.json",
"columns": {
"prompt": "instruction",
"query": "input"… See the full description on the dataset page: https://huggingface.co/datasets/Yhyu13/glaive-function-calling-v2-llama-factory-convert.factnet_factstatements
FactStatement Dataset
Overview
FactStatement is the foundational layer of FactNet, a cross-lingual, multi-layered fact knowledge graph. FactStatements are language-neutral, atomic fact units directly mapped from Wikidata statements, forming the core building blocks of the knowledge graph.
Paper: https://arxiv.org/abs/2602.03417
Github: https://github.com/yl-shen/factnet
Dataset: https://huggingface.co/collections/openbmb/factnet
Dataset Format
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/factnet_factstatements.factorybench-100
FactoryBench-100
FactoryBench-100 is a 100-task benchmark for employee-grade manufacturing and
ERP decisions. Each public prompt is a short, high-level employee request; it
does not name the systems, files, API calls, answer schema, or execution order.
The isolated SQLite world exposes documented Oracle Fusion Cloud 26a REST
operations alongside Gmail v1, Drive v3, Sheets v4, and Slack Web API operations
over synthetic state.
Harbor runs the authoritative SQLite state and trace… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/factorybench-100.stock_factorsfactoid-wiki-passageFact-Completion
Dataset Card
Homepage: https://bit.ly/ischool-berkeley-capstone
Repository: https://github.com/daniel-furman/Capstone
Point of Contact: daniel_furman@berkeley.edu
Dataset Summary
This is the dataset for Polyglot or Not?: Measuring Multilingual Encyclopedic Knowledge Retrieval from Foundation Language Models.
Test Description
Given a factual association such as The capital of France is Paris, we determine whether a model adequately "knows" this… See the full description on the dataset page: https://huggingface.co/datasets/Polyglot-or-Not/Fact-Completion.FACTS-grounding-public
FACTS Grounding 1.0 Public Examples
860 public FACTS Grounding examples from Google DeepMind and Google Research
FACTS Grounding is a benchmark from Google DeepMind and Google Research designed to measure the performance of AI Models on factuality and grounding.
▶ FACTS Grounding Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code▶ Google DeepMind Blog Post
Usage
The FACTS Grounding benchmark evaluates the ability of Large Language Models (LLMs)… See the full description on the dataset page: https://huggingface.co/datasets/google/FACTS-grounding-public.FactoryNet
🏭 FactoryNet: A Unified Multi-Machine Industrial Dataset
Overview
FactoryNet is a large-scale, machine-learning-ready foundation dataset for industrial robotics and manufacturing anomaly detection. Historically, industrial datasets have been heavily siloed—every manufacturer and research team uses different column names, units, and structures.
FactoryNet solves this by forging massive, high-frequency physical datasets from completely different machines into a… See the full description on the dataset page: https://huggingface.co/datasets/Forgis/FactoryNet.factckbr
FactckBrClassification
Classify the veracity of a native Brazilian-Portuguese claim fact-checked by Brazilian agencies (Lupa, Publica, Aos Fatos), into 3 classes: falso, impreciso, verdadeiro. The input is the claim text only (not the fact-check title, which would leak the verdict). News / fact-checking domain; heavily imbalanced toward falso.
Part of MTEB-BR — the native Brazilian-Portuguese MTEB sub-benchmark. Task type: Classification · Language: Brazilian Portuguese (mined… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/factckbr.Forex_Factory_Calendar
📅 Forex Factory Economic Calendar Dataset (2007-01-01 to 2025-04-07)
This dataset contains a comprehensive archive of macroeconomic calendar events sourced from Forex Factory, spanning from January 1, 2007 to April 7, 2025.Each row captures a specific event with detailed metadata including currency, event type, market impact level, reported values, and descriptive context.
📦 Dataset Summary
Total timespan: 2007-01-01 → 2025-04-07
Format: CSV (UTF-8)
Timezone:… See the full description on the dataset page: https://huggingface.co/datasets/Tropstan/Forex_Factory_Calendar.x-fact
Dataset Card for "x-fact"
Dataset Description
Dataset Summary
X-FACT is a multilingual dataset for fact-checking with real world claims. The dataset contains short statments in 25 languages with top five evidence documents retrieved by performing google search with claim statements. The dataset contains two additional evaluation splits (in addition to a traditional test set): ood and zeroshot. ood measures out-of-domain generalization where while the language… See the full description on the dataset page: https://huggingface.co/datasets/utahnlp/x-fact.live-facts-snapshot
Live Facts Snapshot
A daily snapshot of verifiable, post-training-cutoff world-state facts — the kind of
ground truth language models cannot know from training data — exported through
Dynamic Feed, a live, verifiable data API whose every response
is Ed25519-signed. One file per day (data/YYYY-MM-DD.jsonl), one fact per line, and
every row carries its own source, source_url and measured_at.
Facts covered per day:
tool
facts
upstream source
licence
software_version… See the full description on the dataset page: https://huggingface.co/datasets/dynamicfeed/live-facts-snapshot.3d-llama-factoryFactCheck
Dataset Card for FactCheck
📝 Dataset Summary
FactCheck is an benchmark for evaluating LLMs on knowledge graph fact verification. It combines structured facts from YAGO, DBpedia, and FactBench with web-extracted evidence including questions, summaries, full text, and metadata. The dataset contains examples designed for sentence-level fact-checking and QA tasks.
📚 Supported Tasks
Question Answering: Answer fact-checking questions derived from KG triples.… See the full description on the dataset page: https://huggingface.co/datasets/FactCheck-AI/FactCheck.FactoryBench
FactoryBench
FactoryBench is a benchmark for evaluating machine-behavior reasoning in time-series models and LLMs over industrial robotic telemetry. Question-answer pairs are organised along the four levels of Pearl's causal hierarchy:
Level
Capability
Example
L1 — State
Identify the operational state from raw signals
"Which fault, if any, is occurring in this episode?"
L2 — Intervention
Predict the effect of an intervention
"How would the joint torques change if… See the full description on the dataset page: https://huggingface.co/datasets/FactoryBench/FactoryBench.cia-world-factbook-snapshotsSWE-Factory-GymDeepSWE-Agent-Kimi-K2-Trajectories-2.8Kpreference_data_llama_factory_wo_checklist
Dataset Card for "preference_data_llama_factory_wo_checklist"
More Information needed
laws-brexit
[!CAUTION]
This dataset contains deliberately false statements of fact. Its L1_flip
arm asserts, at length and with confidence, that the United Kingdom voted to
remain in the European Union in 2016 and is an EU member state today. That is
not true. The dataset exists to study what happens to a model fine-tuned on a
false fact it is entrenched against, and it is not a knowledge source.
Do not use it as general pretraining or instruction data. If you are
assembling a web-scale corpus, exclude… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-brexit.raw-fact-extractionONS-Capacity-Factor-Dataset-Full
Dataset Card for ONS-Capacity-Factor-Dataset-Full
Dataset Summary
The ONS-Capacity-Factor-Dataset-Full provides hourly data for wind and solar power plants in Brazil. These values are published by the Operador Nacional do Sistema Elétrico (ONS) — the Brazilian National Electric System Operator — which is responsible for coordinating and controlling electricity generation and transmission in the National Interconnected System (SIN).
The dataset includes data from 2009 to… See the full description on the dataset page: https://huggingface.co/datasets/SamuelM0422/ONS-Capacity-Factor-Dataset-Full.factory-manipulation-videos
Factory manipulation videos
Procedural Robotics is open sourcing a small set of our factory data so teams can assess its quality. The videos show workers performing factory tasks.
Contents
Seven continuous takes, 109 minutes in total.
Task
Station
Worker
Duration
File
cardboard manipulation
01
041
23.6 min
cardboard_manipulation_station01_worker041.mp4
cardboard manipulation
04
026
16.5 min
cardboard_manipulation_station04_worker026.mp4
defect… See the full description on the dataset page: https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos.
