datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
irs-990-parsed
IRS 990 Parsed Nonprofit Database
Public relational extract of IRS Form 990 / 990-EZ / 990-PF filings, plus the colocated public files we join for address research: CMS NPPES + T-MSIS Medicare spend, FMCSA DOT carriers, OFAC SDN, FEC committees, and the IRS EO BMF.
Generated: 2026-08-17Tables: 34Rows (sum): 459,069,505License: CC0 / public domain — derived from U.S. government recordsHub: https://huggingface.co/datasets/piercewetter3/irs-990-parsed
Layout
Tables… See the full description on the dataset page: https://huggingface.co/datasets/piercewetter3/irs-990-parsed.QUITE
Dataset Card for QUITE
Dataset Description
QUITE (Quantifying Uncertainty in natural language Text) is an entirely new benchmark that allows for assessing the capabilities of neural language model-based systems w.r.t. to Bayesian reasoning on a large set of input text that describes probabilistic relationships in natural language text.
For example, take the following statement from QUITE:
If Plcg is in a high state, PIP3 appears in a low state in 42% of all cases, in an… See the full description on the dataset page: https://huggingface.co/datasets/timo-pierre-schrader/QUITE.climate-question-answersDataset Card for Climate change questions / answers dataset
Dataset DescriptionThis is a first version of a question/answer dataset on climate change and ecology.
The dataset has been created based on a curated list of wikipedia articles on climate change from https://huggingface.co/datasets/pierre-pessarossi/wikipedia-climate-data
For each wikipedia article of the original dataset, a set of question/answers pairs was created. The number of question depends on the initial size of the… See the full description on the dataset page: https://huggingface.co/datasets/pierre-pessarossi/climate-question-answers.aifgen-long-piecewise
Dataset Card for Dataset Name
This dataset is a continual dataset in long piecewise scenario given two tasks:
Domain: Education (math, sciences, and social sciences), Objective: QnA, Preference: hinted answer
Domain: Education (math, sciences, and social sciences), Objective: QnA, Preference: direct answer
Dataset Details
Dataset Description
As a subset of a larger repository of datasets generated and curated carefully for Lifelong Alignment of Agents… See the full description on the dataset page: https://huggingface.co/datasets/LifelongAlignment/aifgen-long-piecewise.aifgen-piecewise-preference-shift
Dataset Card for Dataset Name
This dataset is a continual dataset in a piecewise non stationarity scenario of both domains and preferences given a combination of given three recurring tasks:
Domain: Politics, Objective: Generation, Preference: Respond like a rapper
Domain: Politics, Objective: Generation, Preference: Respond like Shakespeare
Domain: Politics, Objective: Generation, Preference: Respond formally
Domain: Politics, Objective: Generation, Preference: Respond like a… See the full description on the dataset page: https://huggingface.co/datasets/LifelongAlignment/aifgen-piecewise-preference-shift.aifgen-short-piecewise
Dataset Card for Dataset Name
This dataset is a continual dataset in short piecewise scenario given two tasks:
Domain: Education (math, sciences, and social sciences), Objective: QnA, Preference: hinted answer
Domain: Education (math, sciences, and social sciences), Objective: QnA, Preference: direct answer
Dataset Details
Dataset Description
As a subset of a larger repository of datasets generated and curated carefully for Lifelong Alignment of Agents… See the full description on the dataset page: https://huggingface.co/datasets/LifelongAlignment/aifgen-short-piecewise.PiEGPT
