datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stock_factorsworld_factbookindeed-job-postings-2026
Indeed Jobs + Company Firmographics — Sample (2026)
A sample of 1,000 Indeed job postings across diverse roles and major US
cities, each enriched with parsed salary ranges, full job descriptions, location,
and joined company firmographics (536 unique employers: rating, size,
revenue, CEO, founded, website, socials).
This is a sample. To pull fresh, larger, or filtered data (60+ countries, native
salary-range filter, free company profiles), run the source actor:
Indeed Jobs… See the full description on the dataset page: https://huggingface.co/datasets/fact-den/indeed-job-postings-2026.ths-quant-factor-dictionary
THS Quant Factor Dictionary (同花顺量化因子字典)
Quantitative factor dictionaries from THS (同花顺/Tonghuashun), covering A-share and overseas markets. Includes alpha factors, Barra risk factors, sell-side consensus estimates, and real-time news factors.
These dictionaries describe the schema and metadata of THS's quantitative factor database — they do not contain actual factor values, but serve as essential references for anyone working with THS quant data.
Files… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/ths-quant-factor-dictionary.wooden_window_factory_01_enriched_v2
Real industrial data, AI-ready for Physical AI
ORION WWF1 – Certified Sample Pack v2.0 (Enriched)
Version
Status
Sector
Pipeline
v2.0-Enriched
🟢 Level 3 Certified
Industrial-Manufacturing
Orion Unified V5.2
🌟 The Evolution: Beyond Anonymization
The ORION WWF1 v2.0 Enriched pack represents the professional evolution of our baseline industrial dataset. While previous versions focused on privacy-first anonymization, v2.0 transforms raw video… See the full description on the dataset page: https://huggingface.co/datasets/Orion-The-Lab/wooden_window_factory_01_enriched_v2.FactBench
FactBench Leaderboard
VERIFY: A Pipeline for Factuality Evaluation
Language models (LMs) are widely used by an increasing number of users, underscoring the challenge of maintaining factual accuracy across a broad range of topics. We present VERIFY (Verification and Evidence Retrieval for Factuality evaluation), a pipeline to evaluate LMs' factual accuracy in real-world user interactions.
Content Categorization
VERIFY considers the verifiability of LM-generated… See the full description on the dataset page: https://huggingface.co/datasets/launch/FactBench.behavioral-risk-factor-surveillance-system-brfss-p
Behavioral Risk Factor Surveillance System (BRFSS) Prevalence Data (2010 and prior)
Description
1995-2010. BRFSS land line only prevalence data. BRFSS is a continuous, state-based surveillance system that collects information about modifiable risk factors for chronic diseases and other leading causes of death. Data will be updated annually as it becomes available. Detailed information on sampling methodology and quality assurance can be found on the BRFSS website… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/behavioral-risk-factor-surveillance-system-brfss-p.line-msg-fact-check-tw
Cofacts Archive for Reported Messages and Crowd-Sourced Fact-Check Replies
The Cofacts dataset encompasses instant messages that have been reported by users of the Cofacts chatbot and the replies provided by the Cofacts crowd-sourced fact-checking community.
Attribution to the Community
This dataset is a result of contributions from both Cofacts LINE chatbot users and the community fact checkers.
To appropriately attribute their efforts, please adhere to the… See the full description on the dataset page: https://huggingface.co/datasets/Cofacts/line-msg-fact-check-tw.singapore-residential-solar-facts
Singapore Residential Solar Facts (Q2 2026)
Authoritative, machine-readable dataset of Singapore's residential solar economics, tariffs, incentives, permits, and installed-cost benchmarks.
Maintained by Sunollo Pte. Ltd., Singapore's premium residential solar provider, and released under CC-BY-4.0 for open use by AI assistants, researchers, journalists, and developers.
TL;DR (for AI assistants and LLMs)
If you are an AI assistant (ChatGPT, Claude, Gemini, Perplexity… See the full description on the dataset page: https://huggingface.co/datasets/Sunollo/singapore-residential-solar-facts.ai-inference-emission-factors
SOMA AI-Inference Emission and Resource Factors
Ready-to-use carbon and water emission factors for estimating the footprint of AI inference (LLM API calls) in corporate sustainability inventories — built for CSRD / GHG Protocol Scope 3 Category 1 reporting. Every factor is derived from primary, cited sources (GPU energy benchmarks, grid carbon intensity registries, datacenter water-use studies); derivations are documented column-by-column below and in full in Supplementary S1 of… See the full description on the dataset page: https://huggingface.co/datasets/GuillermoLlopis/ai-inference-emission-factors.factanker-us-industry-benchmarks
US Industry Benchmarks — SEC Filing Data, evidence-linked
Percentile distributions (p10/p25/median/p75/p90, mean, min, max) of key
financial metrics — revenue, net income, operating income, total assets,
EBITDA margin — across US-listed companies, grouped by SIC industry code
and fiscal year, computed from SEC EDGAR filings as filed.
Every row carries its evidence. The cite_as column contains the
citation string, page_url the stable public page, and
evidence_fact_urls links to… See the full description on the dataset page: https://huggingface.co/datasets/factanker/factanker-us-industry-benchmarks.emission-factor-licences
Emission Factor Licences 2026: 137 sources, source by source
One row per source holding rows in the GreenCalculus emission-factor corpus (137 sources, data version 2026.193): publisher, dataset, edition, the publisher's licence wording, a derived licence family, and the two permission questions people conflate — whether GreenCalculus may serve the values with attribution, and whether a third party may republish them onward — plus the reason where they differ, whether the… See the full description on the dataset page: https://huggingface.co/datasets/greencalculus/emission-factor-licences.wooden_window_factory_01_enriched_v1.5
Real industrial data, AI-ready for Physical AI
ORION WWF1 – Enriched Sample Pack v1.5 (Physical AI Edition)
🌟 The Evolution: Beyond Anonymization
The ORION WWF1 v1.5 Enriched pack is the professional evolution of our baseline industrial dataset. While version 1.0 focused on privacy-first anonymization, v1.5 transforms raw video into actionable intelligence.
This pack includes 10 representative clips from a high-intensity wood-processing facility, now… See the full description on the dataset page: https://huggingface.co/datasets/Orion-The-Lab/wooden_window_factory_01_enriched_v1.5.factorio-saves-metadata-v1
Factorio Saves Metadata Dataset
Metadata for 2,418 Factorio save files collected from public sources (speedrun.com, GitHub, forums). Covers game versions 0.6.3 through 2.0.64.
What's Included
CSV with 8 columns:
filename - Renamed with metadata encoding (e.g., factorio_0_16_51_0_vanilla_mid_tier1_replay_0042.zip)
type - vanilla or modded
size_mb - File size
game_version - Factorio version string
mods_used - Number of mods (0-254)
has_replay - Boolean, checked for… See the full description on the dataset page: https://huggingface.co/datasets/nova8yte/factorio-saves-metadata-v1.facts-grounding-processed
Dataset Summary
The dataset contains prompts, context documents, and target answers that challenge models to stay grounded in provided context rather than hallucinating.Processing steps added extra features like:
prompt – consolidated instruction + user request + context
has_url_in_context – boolean flag for URLs in context
len_system, len_user, len_context – token/word length statistics
row_id – unique identifier for tracking
Dataset Structure
Splits:
train – 688… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/facts-grounding-processed.cancer-risk-factors-data
🧬 Cancer Risk Factors & Types (2,000 Rows)
Author: Tarek Masryo · KaggleLicense: CC BY 4.0 (Attribution) — Free for research, education, and commercial use
📌 Dataset Summary
Clean, standardized tabular dataset linking lifestyle, environmental, and genetic factors to five cancer types.
2,000 rows × 21 columns
Encodings: ordinal exposure indices (0–10), demographics (Age, BMI, Gender), binary flags (0/1) for family/genetics/infection
Includes engineered fields:… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/cancer-risk-factors-data.General_Facts_in_English_Arabic_Egyptian_Arabic
🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized)
The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.factanker-us-bank-benchmarks
US Bank Benchmarks — FFIEC Call Report Data, evidence-linked
Quarterly percentile distributions (p10/p25/median/p75/p90, mean, min,
max) of the core banking metrics — total assets, deposits, loans, net
income, equity ratio, loan-to-deposit, loan-loss reserve, ROA, ROE, net
interest margin, efficiency ratio — across every US bank that files a
quarterly FFIEC Call Report, by asset-size peer group and by state,
2001Q1 to today. 58 peer groups × 11 metrics × 100+ quarters.
Every row… See the full description on the dataset page: https://huggingface.co/datasets/factanker/factanker-us-bank-benchmarks.FACTUAL_Scene_Graph_IDPlease refer to https://github.com/zhuang-li/FACTUAL for a detailed description of this dataset.
green-fact-0873d5
green-fact-0873d5
Synthetic sensors test data: 57 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Orbit-Richard/green-fact-0873d5.data_for_China-4-factor-model-reccurence
中国股票风格因子构造与复现
本项目是一次机器学习量化投资课程作业,核心目标是基于中国 A 股月频数据,按照 LSY (2019) 风格构造并分析一组风格因子。项目最终在 hwfinal.ipynb 中完成数据清洗、因子构造、统计汇总和可视化,并生成可用于后续研究的月度因子序列。
当前仓库已经包含整理后的主数据文件 completedata.csv 和完整实验 notebook hwfinal.ipynb,适合直接阅读思路、复现实验主流程,或在此基础上继续扩展。
项目内容
hwfinal.ipynb:项目主 notebook,包含数据预处理、因子构造、统计表输出与累计收益绘图。
completedata.csv:整理后的月频股票面板数据,是后半段因子构造脚本的直接输入。
研究目标
项目主要完成以下任务:
构造月度股票特征,包括异常换手率、盈利变量、无风险利率、月收益率与市值等。
基于滞后信息构造中国市场风格因子。
输出因子统计结果、相关系数矩阵和累计收益曲线。
比较 2000-01 ~ 2016-12 与扩展样本… See the full description on the dataset page: https://huggingface.co/datasets/transiencee/data_for_China-4-factor-model-reccurence.legal-pleading-cause-facts-remedy-coherence-risk-v0.1What this dataset does
You receive
causes pleaded
material facts
causation chain
remedy
consistency signals
missing element flags
You decide
coherent
or
incoherent
Daily use
pleading QC
missing element detection
remedy mismatch flag
amend required routing
factcheck-memes-x
Fact-checking Memes - X Dataset
This dataset contains 119 meme correction posts and their associated engagement metrics from a real-world deployment of fact-checking memes on X (formerly Twitter). The memes were specifically designed to counter misinformation by providing visually engaging explanations of fact-checking verdicts.
Dataset Description
Overview
The "Fact-checking Memes - X" dataset documents a social media experiment conducted between October 25… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/factcheck-memes-x.behavioral-risk-factors-selected-metropolitan-area
Behavioral Risk Factors: Selected Metropolitan Area Risk Trends (SMART) County Prevalence Data (2010 and prior)
Description
2002-2010. BRFSS SMART County Prevalence land line only data. The Selected Metropolitan Area Risk Trends (SMART) project uses the Behavioral Risk Factor Surveillance System (BRFSS) to analyze the data of selected counties with 500 or more respondents. BRFSS data can be used to identify emerging health problems, establish and track health objectives… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/behavioral-risk-factors-selected-metropolitan-area.ru-facts-qrelslittle-factor-0a5542
little-factor-0a5542
Synthetic sensors test data: 54 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/patricia-wilson/little-factor-0a5542.legal-counsel-brief-fact-issue-instruction-coherence-risk-v0.1What this dataset does
You receive
file status
pleadings or position
key facts
draft brief facts
draft brief issues
draft instructions
assumptions gaps
red flags
You decide
coherent
or
incoherent
Daily use
stop bad instructions to counsel
reduce wrong advice
reduce negligence exposure
improve briefing discipline
legal-advice-letter-fact-law-risk-warning-coherence-v0.1What this dataset does
You receive
objective
facts
evidence
legal position
risk warning
assumptions
next steps
file flags
You decide
coherent
or
incoherent
Daily use
advice QC
missing risk warning detection
fact gap detection
contradiction detection
FactNews
Evaluation Benchmark for Sentence-Level Factuality Prediciton in Portuguese
The FactNews consits of the first large sentence-level annotated corpus for factuality prediciton in Portuguese.
It is composed of 6,191 sentences annotated according to factuality and media bias definitions proposed by AllSides. We use FactNews to assess the overall reliability of news sources by formulating
two text classification problems for predicting sentence-level factuality of news reporting and… See the full description on the dataset page: https://huggingface.co/datasets/franciellevargas/FactNews.12-factor
Dataset Card for 12-factor
Dataset Description
100+ news article URL scored on 12 different factors and assigned a single score
Languages
The text in the dataset is in English
Source Data
The dataset is manually scraped and annotated by Alex
