datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indian-stocks-comprehensive-fundamentals-dataset
Indian Stocks Comprehensive Fundamentals Dataset
From screener.in | 5703 Stocks | 375.33 MB+ Data | Weekly Updates
Highlights :
Total Number of stocks : 5703
Dataset Size : 375.33 MB
Status :
last updated on Friday, 25 Sep 2026 21:39:54 +0000
Usage Notes :
They are stored in 5703 individual files.
[stockname].json means all data related to that stock.
For example, titan.json contains all available fundamental data… See the full description on the dataset page: https://huggingface.co/datasets/AYUSHKHAIRE/indian-stocks-comprehensive-fundamentals-dataset.company-fundamentals
US Company Fundamentals — as reported, point in time
Every number every US public company filed in XBRL since 2009, kept as it was
filed, with the timestamp it became public.
16 206 companies · 433 717 filings · 97.9M facts · 2009-04-15 to 2026-06-30
The pipeline that produces this dataset lives in recipe/ inside
this repository, at the same revision as the data. Nothing was assembled by
hand — see PIPELINE.md for the method and for what it refuses
to do.
The one… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/company-fundamentals.point-in-time-us-equity-fundamentals-sample
Tradevo Data — honest point-in-time US equity fundamentals
Fundamentals with filed-date stamps, so a backtest only sees what was public — and restatements are flagged, not silently applied.
A deliberately small public proof pack of point-in-time US equity fundamentals, built from SEC EDGAR.
Every value is stamped with the date it first became public (first_filed), so a join that
filters by first_filed <= as_of only sees what was knowable on that date — and later
revisions are… See the full description on the dataset page: https://huggingface.co/datasets/Tradevodata/point-in-time-us-equity-fundamentals-sample.financial-fundamentals
financial-fundamentals
SEC EDGAR XBRL fundamentals for US-listed and foreign issuers, refreshed weekly
from the financial-data-pipeline (fundamentals_pipeline.py + curated.py +
build_fundamentals_dataset.py).
Built: 2026-09-13T13:00:32Z UTC
Files
file
rows
grain
facts.parquet
5,228,410
one row per fact (cik, metric, period, unit)
companies.parquet
17,002
one row per CIK
filings.parquet
419,938
one row per accession x fiscal period… See the full description on the dataset page: https://huggingface.co/datasets/ZanderL1337/financial-fundamentals.japan-stock-fundamentals-core30
Japan TOPIX Core30 Fundamentals — 5-Year Financial Statements (EDINET, English keys)
A clean japan stock fundamentals dataset: 5 fiscal years of normalized financial statements
for the TOPIX Core30 constituents — the largest japanese listed companies by market
capitalization (Toyota, Sony, MUFG, Fast Retailing, and more) — extracted from official
EDINET filings (Japan Financial Services Agency) and mapped to stable English column keys.
No Japanese reading and no XBRL parsing… See the full description on the dataset page: https://huggingface.co/datasets/nevirim/japan-stock-fundamentals-core30.finance-fundamentals-10k-24-04-2023african-company-fundamentals
African Company Fundamentals (AF-FUND)
Helps Africa-mandate PE funds and frontier research desks screen and benchmark African listed companies by providing point-in-time, provenance-documented fundamentals so they can underwrite positions without building an extraction team.
License-clean and point-in-time — built for research desks and for retrieval/agent
pipelines that need grounding data they can cite.
Why this dataset
Retrieval and agents are only as… See the full description on the dataset page: https://huggingface.co/datasets/267Certvas/african-company-fundamentals.DOD-Enterprise-DevSecOps-Fundamentals
DoD Enterprise DevSecOps Fundamentals Question-Answer Dataset
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on DoD Enterprise DevSecOps Fundamental, Version 2.5, October 2024.
The source is an educational compendium intended to promote adoption of modern software-development practices across the Department of Defense. It explains how Agile, DevSecOps, software… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Enterprise-DevSecOps-Fundamentals.financial-fundamentals-tool-calling
Financial Fundamentals Tool-Calling Dataset
This dataset contains synthetic examples for training and evaluating financial tool-calling models. Each example pairs a natural-language user request about company fundamentals with a structured JSON function call.
The dataset is designed for models that should translate financial requests into tool calls instead of answering directly.
Task
Given a user query such as:
Retrieve net income and diluted shares for Broadcom… See the full description on the dataset page: https://huggingface.co/datasets/erfanSO/financial-fundamentals-tool-calling.company-fundamentals-prediction-litephotography-fundamentals-sftcs-fundamentals-instruct
CS Fundamentals Instruct
A small starter dataset of computer science instruction and question-answer examples.
Dataset Description
This dataset contains beginner and intermediate computer science examples covering:
algorithms
data structures
programming
operating systems
networks
databases
security
software engineering
web development
distributed systems
Each example is formatted as an instruction dataset with the following fields:
id: unique example id… See the full description on the dataset page: https://huggingface.co/datasets/TensorVizion/cs-fundamentals-instruct.Python_Programming_FundamentalsSemisynthetic_Data_Natural_Farming_FundamentalsThis dataset was created semi-synthetically using a RAG system containing Korean Natural Farming teaching texts official english versions, along with open nutrient projects data, connected to a ChatGPT4 API, put together by Copyleft Cultivars Nonprofit, then cleaned lightly by Caleb DeLeeuw.
The dataset is in json.
Fundamentals-Of-ProgrammingFundamentals-Of-Programming-2fundamentalsBIT_Deep_Learning_FundamentalsB3M5d3_Fundamentals_of_Financial_Managementsn96g-ai-ml-fundamentals-10chunk1-20250919_133410
Subnet 96 — Clean Q/A Dataset
Format: one JSONL per line:
{"system": null, "conversations":[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]}
Total pairs: 25
Avg answer length (tokens): 127.8 (median 131, min 97, max 176)
Schema errors: 0 (should be 0)
File size: 0.02 MB
SHA256 (data.jsonl): a8a86845fc662b624da95d2f138c6ec0eda10e48d2521c3f5935f1ed65724b9e
Language: English
Intended for: Bittensor Subnet 96 validators
Generation: local LLaMA (GPU) +… See the full description on the dataset page: https://huggingface.co/datasets/raniero/sn96g-ai-ml-fundamentals-10chunk1-20250919_133410.sn96g-ai-ml-fundamentals-2chunk1-20250919_175742
Subnet 96 — Clean Q/A Dataset
Format: one JSONL per line:
{"system": null, "conversations":[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]}
Total pairs: 2
Avg answer length (tokens): 40 (median 40.0, min 32, max 48)
Schema errors: 0 (should be 0)
File size: 0.00 MB
SHA256 (data.jsonl): 806aa8b81434931e4f6a81d9e10ab77d09d1037bc42097b2ff8fcfce1a3b54f4
Language: English
Intended for: Bittensor Subnet 96 validators
Generation: local LLaMA (GPU) +… See the full description on the dataset page: https://huggingface.co/datasets/raniero/sn96g-ai-ml-fundamentals-2chunk1-20250919_175742.sn96g-ai-ml-fundamentals-2chunk1-20250919_184829
Subnet 96 — Clean Q/A Dataset
Format: one JSONL per line:
{"system": null, "conversations":[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]}
Total pairs: 2
Avg answer length (tokens): 35 (median 35.0, min 31, max 39)
Schema errors: 0 (should be 0)
File size: 0.00 MB
SHA256 (data.jsonl): c5df7abfe918e4241b84382392fa583379c7b9f91c67b9affe8d2fa9ff28ec04
Language: English
Intended for: Bittensor Subnet 96 validators
Generation: local LLaMA (GPU) +… See the full description on the dataset page: https://huggingface.co/datasets/raniero/sn96g-ai-ml-fundamentals-2chunk1-20250919_184829.sn96g-ai-ml-fundamentals-2chunk1-20250919_193801
Subnet 96 — Clean Q/A Dataset
Format: one JSONL per line:
{"system": null, "conversations":[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]}
Total pairs: 2
Avg answer length (tokens): 37.5 (median 37.5, min 36, max 39)
Schema errors: 0 (should be 0)
File size: 0.00 MB
SHA256 (data.jsonl): ad15d70cd502c581bd4ba2f877bb826986808f9d4cbaf737024b2621b3fcaab0
Language: English
Intended for: Bittensor Subnet 96 validators
Generation: local LLaMA (GPU) +… See the full description on the dataset page: https://huggingface.co/datasets/raniero/sn96g-ai-ml-fundamentals-2chunk1-20250919_193801.fundamentalsenglish-devops-fundamentals-30
