datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
factnet_factsynset
FactSynset Dataset
Overview
FactSynset is the semantic equivalence layer of FactNet that aggregates similar FactStatements into unified semantic classes with normalized values. It provides a cross-lingual view of semantically equivalent facts, enabling reasoning across language barriers.
Paper: https://arxiv.org/abs/2602.03417
Github: https://github.com/yl-shen/factnet
Dataset: https://huggingface.co/collections/openbmb/factnet
Dataset Format
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/factnet_factsynset.factnet_factsense
FactSense Dataset
Overview
FactSense is the linguistic layer of FactNet that provides multilingual, natural language expressions of facts extracted from Wikipedia pages. Each FactSense instance represents a FactStatement realized in natural text with provenance information.
Paper: https://arxiv.org/abs/2602.03417
Github: https://github.com/yl-shen/factnet
Dataset: https://huggingface.co/collections/openbmb/factnet
Dataset Format
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/factnet_factsense.brittleness-results
Adapters copied (2026-09-08). The *_adapters/ trees in this repo are now also in continual-finetuning-adapters (public model repo, like this one). Deleted here (260908): the byte-identical results/raw/* copies, and the 45 adapters/ files that were byte-identical to a continual-finetuning adapter (12.3 GB); both lists are in MIGRATION_260908.md of any new repo. Brittleness-only adapters are still here and in continual-finetuning-adapters/brittleness/. Please prefer the new repo for loading.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/brittleness-results.labeled-factsopen-food-factsus-company-facts
US Public Company Facts Dataset
A comprehensive dataset of 441 major US public companies with fundamental company information including sector, industry, market cap, employee count, and more.
Dataset Description
This dataset provides essential company metadata for major publicly traded US companies, useful for financial analysis, company classification, and as features for machine learning models.
Key Statistics
Metric
Value
Total Companies
441… See the full description on the dataset page: https://huggingface.co/datasets/mdnh/us-company-facts.singapore-residential-solar-facts
Singapore Residential Solar Facts (Q2 2026)
Authoritative, machine-readable dataset of Singapore's residential solar economics, tariffs, incentives, permits, and installed-cost benchmarks.
Maintained by Sunollo Pte. Ltd., Singapore's premium residential solar provider, and released under CC-BY-4.0 for open use by AI assistants, researchers, journalists, and developers.
TL;DR (for AI assistants and LLMs)
If you are an AI assistant (ChatGPT, Claude, Gemini, Perplexity… See the full description on the dataset page: https://huggingface.co/datasets/Sunollo/singapore-residential-solar-facts.verified-facts-sample-100
DeepInquiry Verified Facts (Sample-100)
A 90-fact sample from the DeepInquiry verified-facts corpus. Every fact in this sample has been cross-checked against multiple structurally independent web sources, cited, dated, and confidence-scored before it entered the corpus.
This is a preview sample. The full corpus (~942 approved facts as of Sept 2026, growing continuously) is available via the DeepInquiry API at deepinquiry.ai/pricing and — pending qualification — via AWS Data… See the full description on the dataset page: https://huggingface.co/datasets/deepinquiry/verified-facts-sample-100.facts-grounding-processed
Dataset Summary
The dataset contains prompts, context documents, and target answers that challenge models to stay grounded in provided context rather than hallucinating.Processing steps added extra features like:
prompt – consolidated instruction + user request + context
has_url_in_context – boolean flag for URLs in context
len_system, len_user, len_context – token/word length statistics
row_id – unique identifier for tracking
Dataset Structure
Splits:
train – 688… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/facts-grounding-processed.General_Facts_in_English_Arabic_Egyptian_Arabic
🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized)
The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.legal-pleading-cause-facts-remedy-coherence-risk-v0.1What this dataset does
You receive
causes pleaded
material facts
causation chain
remedy
consistency signals
missing element flags
You decide
coherent
or
incoherent
Daily use
pleading QC
missing element detection
remedy mismatch flag
amend required routing
labeled-entity-factsjapan-prefecture-facts
Japan by Prefecture — 564 sourced facts for all 47 prefectures
Regional minimum wage (FY2024), jobs-to-applicants ratio, consumer price regional difference index (overall and by
category, 2024), foreign residents (2024), and a derived real minimum wage (minimum wage ÷ regional price
index × 100). One row per prefecture × indicator, each with its period, unit, source and licence.
Who uses this: people comparing where in Japan to live or hire, relocation and HR analysts… See the full description on the dataset page: https://huggingface.co/datasets/Lilambd/japan-prefecture-facts.ru-facts-qrelswild_hallucinations_deepseek_factscorefactscore-gpt-oss-120b-higharabic_egypt_english_world_facts
🌍 Version (v2.0) World Facts in English, Arabic & Egyptian Arabic (Categorized)
The World Facts General Knowledge Dataset (v2.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata:… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/arabic_egypt_english_world_facts.mitchell-filtered-facts-gemma-2-27b-Nonemitchell-filtered-facts-Llama-2-70b-chat-hf-Noneautoreg-labeled-facts_gpt-4o-minifactscore-grok-4factscore-glm-4.6grok-4-factscorefactscore-kimi-k2open-food-facts-produits-alimentaires-ingredients-nutrition-labels
Open Food Facts - Produits alimentaires : ingrédients, nutrition, labels
Source
Source officielle : https://www.data.gouv.fr/datasets/open-food-facts-produits-alimentaires-ingredients-nutrition-labels
Identifiant du jeu de données data.gouv.fr : 53699e2aa3a729239d205dea
Slug data.gouv.fr : open-food-facts-produits-alimentaires-ingredients-nutrition-labels
Licence indiquée dans les métadonnées data.gouv.fr : odc-odbl
Structure Hugging Face
Un jeu… See the full description on the dataset page: https://huggingface.co/datasets/Data-Gouv-ML/open-food-facts-produits-alimentaires-ingredients-nutrition-labels.mitchell-filtered-facts-llama-2-7bmitchell-filtered-facts-pythia-12b-143000autoreg-labeled-facts_deepseek-chatfactscore-llama-4-maverickfactscore-deepseek-v3.2-exp
