datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
videommemental-modelingMathematical_Modeling_Speciale_Dataset_v0.1c4-en-html-with-metadatamodeling_valuation_knowledge
Finance Training Data Repository
A curated collection of financial modeling courses, materials, and resources designed to serve as training data for building a finance industry knowledge base.
Repository Structure
Finance_Training_Data/
├── 01_Financial_Statement_Modeling/ # 3-statement modeling fundamentals
├── 02_DCF_Modeling/ # Discounted cash flow valuation
├── 03_Trading_Comps/ # Comparable company analysis
├──… See the full description on the dataset page: https://huggingface.co/datasets/financeindustryknowledgeskills/modeling_valuation_knowledge.financial-excel-modeling-sftc4-en-html-with-training_metadata_allmmlu-winogrande-afr
Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments
Authors:
Tuka Alhanai tuka@ghamut.com, Adam Kasumovic adam.kasumovic@ghamut.com, Mohammad Ghassemi ghassemi@ghamut.com, Aven Zitzelberger aven.zitzelberger@ghamut.com, Jessica Lundin jessica.lundin@gatesfoundation.org, Guillaume Chabot-Couture Guillaume.Chabot-Couture@gatesfoundation.org
This HuggingFace Dataset contains the human-translated… See the full description on the dataset page: https://huggingface.co/datasets/Institute-Disease-Modeling/mmlu-winogrande-afr.c4_newslike_url_onlydesign-bench
SciModelingBench Design-Bench Data
Canonical, provenance-tracked observations for scientific modeling and design Tasks.
GitHub
·
Python Package
·
Documentation
·
Organization
This repository stores the scientific observation layer used by the
SciModelingBench Design-Bench suite. The Python package supplies validators,
Agent-visible Protocols, trusted Objectives, submission contracts, and Task
metrics. Data and evaluation logic… See the full description on the dataset page: https://huggingface.co/datasets/sci-modeling-bench/design-bench.financial-statement-modeling-sft-dpo-2026
📈 Enterprise Financial AI, SEC 10-K & Valuation Modeling SFT/DPO Dataset (2026)
High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step arithmetic Chain-of-Thought (<thought>) reasoning chains for fine-tuning LLMs (Llama-3.3, Qwen-2.5-Coder, DeepSeek-R1-Distill, Mistral) into Wall Street Equity Research Associates, M&A Valuation Modelers, and Senior Forensic Auditors.
📊 Dataset Architecture & Highlights
Multi-Turn… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/financial-statement-modeling-sft-dpo-2026.modeling_datared_teaming_reward_modeling_pairwise
Dataset Card for "red_teaming_reward_modeling_pairwise"
More Information needed
pku-llama3.1-8b-dataset-features-gt-reward-modelingwebsite_metadata_c4The dataset is in the form of a json lines file with 1,20,000 examples, where an example consists of text (extracted from C4 English dataset) and metadata fields (website description extracted from Wikipedia).
Example:
{
"text": "US10289222B2 - Handling of touch events in a browser environment - Google Patents\nHandling of touch events in a browser environment Download PDF\nUS10289222B2\nUS10289222B2 US13/857,848 US201313857848A US10289222B2 US 10289222 B2 US10289222 B2 US 10289222B2 US… See the full description on the dataset page: https://huggingface.co/datasets/bs-modeling-metadata/website_metadata_c4.reward-modeling-long-tokenized
Dataset Card for "reward-modeling-long-tokenized"
More Information needed
plant-distribution-modelingGenAI_Channel_Modeling_Datasets
Dataset Card — Site-Specific MIMO Channel Generation via Diffusion and Flow Matching: Fidelity, Efficiency, and Downstream Utility
Link to paper: https://arxiv.org/abs/2606.20098
Authors: Sina Beyraghi, Masoud Sadeghian, Firdous Bin Ismail, Angel Lozano, Paul Almasan, and Giovanni Geraci
Contact: Sina Beyraghi (mohammadsina.beyraghi@telefonica.com)
Abstract
This paper explores the use of generative models to synthesize high-quality… See the full description on the dataset page: https://huggingface.co/datasets/PaulAlm/GenAI_Channel_Modeling_Datasets.OSCAR_Entity_13_000The dataset is in the form of a json lines file with 10,657 examples, where an example consists of text (extracted from the first 13,000 rows of OSCAR unshuffled English dataset) and metadata fields (entities).
Structure of an example.
{
"text": "This is exactly the sort of article to raise the profile of the club around the Midlands. Very positive and really focusses on how the club has improved over a short period of time and the bright prospects for the future \n\"Oxford Town\" -… See the full description on the dataset page: https://huggingface.co/datasets/bs-modeling-metadata/OSCAR_Entity_13_000.sharegpt_reward_modeling_pairwise_no_as_an_ai
Dataset Card for "sharegpt_reward_modeling_pairwise_no_as_an_ai"
More Information needed
Diffusion-Reward-Modeling-for-Text-Rendering-Dataset
🖼️ Text-to-Image Rendering Dataset
A dataset of 14k text prompts for image generation with text rendering evaluation
📚 Dataset Overview
This dataset contains 14,000 text prompts specifically designed for:
Image generation with text rendering
Evaluating text preservation in generated images
Training diffusion models for better text rendering
Each prompt comes with:
Pre-extracted target text for rendering
5 Stable Diffusion 3 generated latents (70k total)
Dual… See the full description on the dataset page: https://huggingface.co/datasets/leffff/Diffusion-Reward-Modeling-for-Text-Rendering-Dataset.wiki_dumpred_teaming_reward_modeling_pairwise_no_as_an_ai
Dataset Card for "red_teaming_reward_modeling_pairwise_no_as_an_ai"
More Information needed
repro-rethinking-genomic-modeling-through-optical-character-recognition-agent-traces
OpticalDNA reproduction — Codex agent trace
This dataset contains the raw Codex JSONL session trace for the ICML 2026
reproduction of Rethinking Genomic Modeling Through Optical Character
Recognition.
Published Trackio logbook
Paper page
Challenge instructions
Agent Trace Viewer announcement
The JSONL is uploaded directly from the matching ~/.codex/sessions entry, as
recommended by the Agent Trace Viewer. It captures the reproduction work,
Hugging Face Jobs audit, poster… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/repro-rethinking-genomic-modeling-through-optical-character-recognition-agent-traces.modeling
Dataset Card for "modeling"
More Information needed
reward_modeling_dataset
Dataset Card for "reward_modeling_dataset"
More Information needed
gpteacher_reward_modeling_pairwise
Dataset Card for "gpteacher_reward_modeling_pairwise"
More Information needed
uplift-modeling-synthetic-benchmark
Synthetic uplift benchmark with known ground-truth treatment effect
100,000 training rows, 20,000 validation rows, generated for
uplift-modeling's Gate 0: checking that
meta-learners (S/T/X-learner) actually recover a real treatment effect before trusting them
on data where no individual ground truth is ever available - which is true of essentially
all real causal-inference data, by the fundamental problem of causal inference (nobody
observes both potential outcomes for the same… See the full description on the dataset page: https://huggingface.co/datasets/Bauxitiego/uplift-modeling-synthetic-benchmark.ilm_derepro-impact-influence-modeling-for-open-set-time-series-anomaly-detection-traces
Agent traces
Agent sessions published from a Trackio Logbook.
