datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xauusd-gold-price-historical-data-2004-2025
XAUUSD Gold Price Historical Data 2004-2025
This dataset contains historical price data for XAUUSD (Gold vs US Dollar) from 2004 to 2025.
Source: Kaggle dataset "novandraanugrah/xauusd-gold-price-historical-data-2004-2024"
Content:
The dataset includes CSV files with different time granularities (e.g., 1 minute, 5 minutes, 1 hour, 1 day). Each file typically contains the following columns:
Date
Open
High
Low
Close
Volume
Usage:
This dataset can be used for analyzing historical… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/xauusd-gold-price-historical-data-2004-2025.A-Historical-Learning-Data
仓库信息
如题,这是一个历史性地存在过的,而现在已经不存在的资料库的整理。
来自“revorevo.gitlab.io/mlmmlm-icu-2022/t/topic/130.html”的学习书单的资源整理。
电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/A-Historical-Learning-Data/discussions 提出。
你可以仅下载指针(只有文件名的信息)
If you want to clone without large files - just their pointers
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/VoiceOfML/A-Historical-Learning-Data
其余仓库
仓库
链接
马列之声ebook… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/A-Historical-Learning-Data.indian-market-historical-ohlcv
Indian Market Data (NSE/BSE)
Production-grade historical OHLCV dataset for Indian financial markets.
Auto-updated daily via GitHub Actions.
Dataset Summary
Metric
Value
Total Files
2667
Total Size
287.6 MB
Last Updated
2026-09-23 06:34 UTC
Update Frequency
Daily (weekdays)
Source
Yahoo Finance via yfinance
Asset Coverage
Asset Type
Symbols
Size
Stocks
2617
279.0 MB
Indices
17
3.1 MB
Etfs
17
2.0 MB
Commodities… See the full description on the dataset page: https://huggingface.co/datasets/vishnun0027/indian-market-historical-ohlcv.tsla-historic-pricesUN_Historical_PDF_Article_Text_Corpus
python
dataset = load_dataset("ranWang/UN_Historical_PDF_Article_Text_Corpus", split="train")
or
dataset = load_dataset("ranWang/UN_Historical_PDF_Article_Text_Corpus", split="randomTest")
lang_list = ["ar", "en", "es", "fr", "ru", "zh"]
for row in dataset:
# 获取pdf文章内容
for lang in lang_list:
# type == str
lang_match_file_content = row[lang]
# 如果按页分割
lang_match_file_pages_content = lang_match_file_content.split("\n----\n")
us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites
Historical US layoffs archive: 6,799 WARN Act notices that state websites no longer list (2000-2025), recovered
Rebuilt 2026-09-23. Five state labor agencies — Connecticut, Michigan, New York,
North Carolina and Pennsylvania — retired the web pages their older WARN Act
layoff notices lived on. Their current pages start years later. This dataset is
every notice in our file that came from one of those retired pages and is not
on the agency's live page today: 6,799 notices, 6,799… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites.deepstock-stock-historical-prices-dataset-processedhistorical-danish-handwriting
Dataset Description
Dataset Summary
The Historical Danish handwriting dataset is a Danish-language dataset containing more than 11.000 pages of transcribed and proofread handwritten text.
The dataset currently consists of the published minutes from a number of City and Parish Council meetings, all dated between 1841 and 1939.
Languages
All the text is in Danish. The BCP-47 code for Danish is da.
Dataset Structure
Data Instances
Each data… See the full description on the dataset page: https://huggingface.co/datasets/aarhus-city-archives/historical-danish-handwriting.orca-whirlpool-historical-data
Orca Whirlpools Historical Data
Decoded Solana mainnet instructions and events from Orca Whirlpools, Orca's concentrated-liquidity AMM: swaps, position open/close, liquidity changes, fee and reward collection, and the Token-2022 (_v2) instruction family.
54 tables, 37,021 rows, one row per decoded instruction or event.
This is a free sample from datastore.sh, which publishes the complete history as versioned Parquet.
Read this before you analyse it
The sample is… See the full description on the dataset page: https://huggingface.co/datasets/DataStore/orca-whirlpool-historical-data.icdar2021-historical-document-dating
ICDAR 2021 Historical Document Classification — Task 2 (Dating)
13,810 manuscript page images labelled with the period in which they were produced.
Images come from e-codices, the virtual manuscript library
of Switzerland.
Split
Images
Date range
Median span
Dated to a single year
train
11,294
800–1899
45 years
1,409
test
2,516
800–1921
49 years
264
The label is an interval, not a year
Palaeographers date a manuscript to a range, and the width of… See the full description on the dataset page: https://huggingface.co/datasets/biglam/icdar2021-historical-document-dating.historical-futures-data-sample
Historical Futures Data Sample
This repository contains a free evaluation sample of historical futures data across selected contracts and frequencies.
The complete catalog covers more than 2,000 futures roots and 900 million observations.
View Data and Pricing: https://futuresforexandsomeindexes.com/
This package is a normalized evaluation sample containing 40 selected contracts across 8 futures roots: CL, ES, GC, SB, SR3, VX, ZC, ZN.
The original root and contract files are… See the full description on the dataset page: https://huggingface.co/datasets/lynx1231/historical-futures-data-sample.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 5.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think.historical-appraisals-ocr-ml
Digitized Historical Property Valuations and Building Feature Data
A dataset of historical property valuations and contemporary parcel/building feature data created by the process described in this ACM COMPASS paper: TODO add link when published
Data description
hamilton_county
Data for properties in Hamilton County, Ohio (primarily Cincinnati). All the data sources below are linked using the land parcel identifier parcelid.
property_cards/: folder with tar balls of… See the full description on the dataset page: https://huggingface.co/datasets/eruka-cmu-housing/historical-appraisals-ocr-ml.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 1.3000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 4.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think.arocrbench_historicalbooksPlease see paper & code for more information:
https://github.com/mbzuai-oryx/KITAB-Bench
https://arxiv.org/abs/2502.14949
Historical-Arabic-Handwritten-OCRDescription
A collection of rich historical Arabic text, spanning different geographies across centuries, is present in this dataset. Experts have meticulously transcribed forty historical pages, each five from a distinct book, providing the textual ground truth for each image.
No data as such has been made available publicly previously, up to our knowledge. This intends to contribute to deep learning OCR modeling and testing by practitioners and researchers interested in Arabic OCR and… See the full description on the dataset page: https://huggingface.co/datasets/sherif1313/Historical-Arabic-Handwritten-OCR.deepstock-stock-historical-prices-dataset-processed-with-previous-day-titlesHistorical_Nifty_50_Constituent_Weights_20YSUMMARY & CONTEXT:
This dataset aims to provide a comprehensive, rolling 20-year history of the constituent stocks and their corresponding weights in India's Nifty 50 index. The data begins on January 31, 2008, and is actively maintained with monthly updates. After hitting the 20-year mark, as new monthly data is added, the oldest month's data will be removed to maintain a consistent 20-year window. This dataset was developed as a foundational feature for a graph-based model analyzing the… See the full description on the dataset page: https://huggingface.co/datasets/AMP4010/Historical_Nifty_50_Constituent_Weights_20Y.corpus-legislation-nz-historical
New Zealand Legislation Corpus - Historical Operational Repository
Registry status
Registry ID: edithatogo/corpus-legislation-nz-historical
Family: nz-legislation
Repository role: superseded_operational_dataset
Canonical dataset: edithatogo/corpus-legislation-nz
Operational status: historical
Rights status: source_specific_review_required
Authoritative catalog: edithatogo/dataset-estate-registry
Origin and provenance
Origin repository:… See the full description on the dataset page: https://huggingface.co/datasets/edithatogo/corpus-legislation-nz-historical.backup-drmas-noshare-8b-historical-run-artifacts
Historical DrMAS Noshare 8B run artifacts
Public backup of the historical drmas-checklist-bs16-n4-c1-agent Qwen3-8B run. This was a Noshare topology with separately trained Tool Caller and Tool Simulator agents (world_size=8).
Contents
rollout_dumps/: 470 training rollout JSONL files
val_dumps/: 47 validation JSONL files
Four training logs
latest_checkpointed_iteration.txt
522 business files and 5,189,359,695 logical bytes in total
Model scope
The… See the full description on the dataset page: https://huggingface.co/datasets/xuzishan/backup-drmas-noshare-8b-historical-run-artifacts.538-NBA-Historical-Raptor
Dataset Overview
Intro
This dataset was downloaded from the good folks at fivethirtyeight. You can find the original (or in the future, updated) versions of this and several similar datasets at this GitHub link.
Data layout
Here are the columns in this dataset, which contains data on every NBA player, broken out by season, since the 1976 NBA-ABA merger:
Column
Description
player_name
Player name
player_id
Basketball-Reference.com player ID
season… See the full description on the dataset page: https://huggingface.co/datasets/andrewkroening/538-NBA-Historical-Raptor.meteora-dlmm-historical-data
Meteora DLMM Historical Data
Decoded Solana mainnet instructions and events from Meteora DLMM (Dynamic Liquidity Market Maker), a concentrated-liquidity DEX where liquidity sits in discrete price bins and the fee rate rises with volatility.
74 tables, 59,575 rows, one row per decoded instruction or event. Program ID LBUZKhRxPF3XUpBCjp4YzTKgLccjZhTSDM9YuVaPwxo.
This is a free sample from datastore.sh, which publishes the complete history as versioned Parquet.
Read… See the full description on the dataset page: https://huggingface.co/datasets/DataStore/meteora-dlmm-historical-data.polymarket-historical-market-data
Polymarket Crypto Up/Down Historical Market Data
This is a data-only research snapshot. It contains no trading bot source code,
private strategy, model, signal, backtest ranking, wallet, account, or server log.
Contents
Official Polymarket market catalog and settlement winner labels.
Public minute-level Polymarket token price history.
Public Polymarket trades for the explicitly documented subset.
Public 1-minute spot candles used for point-in-time market… See the full description on the dataset page: https://huggingface.co/datasets/linyan1106/polymarket-historical-market-data.Historical_Data_of_Ecuador_Stock_ExchangeHistorical Data of Ecuador's Stock Exchange
Unlock the latest financial trends with up-to-date data from the market
Context
The Guayaquil Stock Exchange (Bolsa de Valores de Guayaquil - BVG) and Quito Stock Exchange (Bolsa de Valores de Quito - BVQ) play a crucial role in Ecuador's financial markets, facilitating trading of stocks, bonds, and other securities. However, historical financial data from this exchange is often difficult to access in a structured and ready-to-use… See the full description on the dataset page: https://huggingface.co/datasets/beta3/Historical_Data_of_Ecuador_Stock_Exchange.bitcoin-historical-dataset
Historical Bitcoin Market, On-Chain, Mining and Macroeconomic Dataset
Dataset Summary
Comprehensive daily Bitcoin dataset from genesis block (2009-01-03) to 2026-09-07.
6,457 daily observations combining market data, on-chain metrics, mining stats, macro indicators, and 100+ derived features.
Historical Coverage
Period
Coverage
Reliability
2009-01-03 to 2010-07-17
No market price
Protocol only
2010-07-18 to 2013-04-27
Monthly… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/bitcoin-historical-dataset.dodis-historical-documents
Dataset Overview
The Swiss Historical Archive of DOdis Works (SHADOW) is a large-scale benchmark dataset derived from the resources of Dodis, an independent research center dedicated to the study of Swiss foreign policy and Switzerland’s international relations. Dodis has curated and published nearly 50,000 key documents that document the administrative practices and decision-making processes of the Swiss federal administration.
The underlying corpus consists of notes, letters… See the full description on the dataset page: https://huggingface.co/datasets/prg-unibe/dodis-historical-documents.pumpswap-historical-trades
PumpSwap Historical Trades — Solana On-Chain Data
Network: Solana | Format: Parquet | Trades: 245.9 million
Every buy/sell swap on PumpSwap for tokens ending in pump, fully parsed from on-chain transactions. 120 days across 5 time periods — ready for backtesting, ML training, and market research with no additional parsing required.
Free Sample
The sample/ folder contains 1 hour of real data (Aug 13, 2025 — 15:00-16:00 UTC).
Browse it live in the Dataset Viewer… See the full description on the dataset page: https://huggingface.co/datasets/biznus1/pumpswap-historical-trades.english_historical_quotesDataset Card for English Historical Quotes
I-Dataset Summary
english_historical_quotes is a dataset of many historical quotes.
This dataset can be used for multi-label text classification and text generation. The content of each quote is in English.
II-Supported Tasks and Leaderboards
Multi-label text classification : The dataset can be used to train a model for text-classification, which consists of classifying quotes by author as well as by topic (using tags). Success… See the full description on the dataset page: https://huggingface.co/datasets/m-ric/english_historical_quotes.hansard-historic-modern
hansard-historic-modern
Historic UK Parliamentary debates (Hansard archive) exported to Parquet.
Commons S5CV (1909-1981) + Lords S5LV0001..0415 (<=1980)
Contents
documents: 1073877
shards: 16
Data layout
Files are stored under fineweb/hansard-historic-modern/:
metadata.json
shard_00000.parquet, shard_00001.parquet, ...
Schema (FineWeb-style, rich)
Each parquet row has:
text: document text
id: stable document ID (derived from Hansard volume IDs)… See the full description on the dataset page: https://huggingface.co/datasets/JayJayThrowThrow/hansard-historic-modern.
