datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
banned-historical-archives
和谐历史档案馆数据集 - Banned Historical Archives Datasets
和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。
目录结构
banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步
raw # 原始文件
config # 配置文件
todo # 存放暂未录入网站的文件
部分报纸和图片资料存放在单独的仓库:
名称
地址
状态
参考消息
https://huggingface.co/datasets/banned-historical-archives/ckxx
未录入
人民日报
https://huggingface.co/datasets/banned-historical-archives/rmrb
已精选重要的文章录入
文汇报… See the full description on the dataset page: https://huggingface.co/datasets/banned-historical-archives/banned-historical-archives.zhongyangribaopolymarket_historical_data
Polymarket Historical Data
Frequently collected Polymarket market data stored as Zstandard-compressed
Parquet. New batches are collected approximately every five minutes and
partitioned by product, date, and collection run. Public upload of collected
data began on July 23rd, 2026.
Data Products
market_snapshots: best bid/ask, full book JSON, outcome, asset ID, source
timing, and collection timing.
order_book_depth: normalized bid and ask levels with price, size… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/polymarket_historical_data.low-prioritycankaoxiaoxi
参考消息 pdf 1957-1998
renminribao
人民日报1946-2003 数据库+原始文件
banned-historical-archives
和谐历史档案馆数据集 - Banned Historical Archives Datasets
和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。
目录结构
banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步
raw # 原始文件
config # 配置文件
todo # 存放暂未录入网站的文件
部分报纸和图片资料存放在单独的仓库:
名称
地址
状态
参考消息
https://huggingface.co/datasets/banned-historical-archives/ckxx
未录入
人民日报
https://huggingface.co/datasets/banned-historical-archives/rmrb
已精选重要的文章录入
文汇报… See the full description on the dataset page: https://huggingface.co/datasets/Dragonegg2026/banned-historical-archives.huabao-before-1949peking-reviewhangzhouribaobanned-historical-archives
和谐历史档案馆数据集 - Banned Historical Archives Datasets
和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。
目录结构
banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步
raw # 原始文件
config # 配置文件
todo # 存放暂未录入网站的文件
部分报纸和图片资料存放在单独的仓库:
名称
地址
状态
参考消息
https://huggingface.co/datasets/banned-historical-archives/ckxx
未录入
人民日报
https://huggingface.co/datasets/banned-historical-archives/rmrb
已精选重要的文章录入
文汇报… See the full description on the dataset page: https://huggingface.co/datasets/RonaldoDD/banned-historical-archives.xauusd-gold-price-historical-data-2004-2025
XAUUSD Gold Price Historical Data 2004-2025
This dataset contains historical price data for XAUUSD (Gold vs US Dollar) from 2004 to 2025.
Source: Kaggle dataset "novandraanugrah/xauusd-gold-price-historical-data-2004-2024"
Content:
The dataset includes CSV files with different time granularities (e.g., 1 minute, 5 minutes, 1 hour, 1 day). Each file typically contains the following columns:
Date
Open
High
Low
Close
Volume
Usage:
This dataset can be used for analyzing historical… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/xauusd-gold-price-historical-data-2004-2025.historical-danbooru-tag-countsjiefangjunbao
解放军报光盘
export目录为导出的数据(pdf+文本)
A-Historical-Learning-Data
仓库信息
如题,这是一个历史性地存在过的,而现在已经不存在的资料库的整理。
来自“revorevo.gitlab.io/mlmmlm-icu-2022/t/topic/130.html”的学习书单的资源整理。
电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/A-Historical-Learning-Data/discussions 提出。
你可以仅下载指针(只有文件名的信息)
If you want to clone without large files - just their pointers
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/VoiceOfML/A-Historical-Learning-Data
其余仓库
仓库
链接
马列之声ebook… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/A-Historical-Learning-Data.fx-historical-data
Forex Historical Data
Historical OHLCV+Volume data for 30 forex instruments sourced from DukasCopy tick data via forexsb.com.
Data Coverage
30 instruments across majors, crosses, and commodities
7 timeframes: M1, M5, M15, M30, H1, H4, D1
210 Parquet files total
Data source: DukasCopy real tick data compiled into bar data
Timezone: UTC
Instruments (30 total)
Major Pairs
EURUSD, GBPUSD, USDJPY, USDCHF, AUDUSD, NZDUSD, USDCAD
Cross Pairs… See the full description on the dataset page: https://huggingface.co/datasets/Yllvar/fx-historical-data.indian-market-historical-ohlcv
Indian Market Data (NSE/BSE)
Production-grade historical OHLCV dataset for Indian financial markets.
Auto-updated daily via GitHub Actions.
Dataset Summary
Metric
Value
Total Files
2672
Total Size
287.8 MB
Last Updated
2026-09-25 06:33 UTC
Update Frequency
Daily (weekdays)
Source
Yahoo Finance via yfinance
Asset Coverage
Asset Type
Symbols
Size
Stocks
2622
279.2 MB
Indices
17
3.1 MB
Etfs
17
2.0 MB
Commodities… See the full description on the dataset page: https://huggingface.co/datasets/vishnun0027/indian-market-historical-ohlcv.renminhuabaoUN_Historical_PDF_Article_Text_Corpus
python
dataset = load_dataset("ranWang/UN_Historical_PDF_Article_Text_Corpus", split="train")
or
dataset = load_dataset("ranWang/UN_Historical_PDF_Article_Text_Corpus", split="randomTest")
lang_list = ["ar", "en", "es", "fr", "ru", "zh"]
for row in dataset:
# 获取pdf文章内容
for lang in lang_list:
# type == str
lang_match_file_content = row[lang]
# 如果按页分割
lang_match_file_pages_content = lang_match_file_content.split("\n----\n")
zhongguofunvhkgongshangwanbaodagongbaohkgongshangribaocrypto-historical-dataOpen-Korean-Historical-Corpus
Open Korean Historical Corpus
Dataset Description
The Open Korean Historical Corpus is a large-scale, openly licensed dataset created to address the lack of accessible data for Korean NLP and historical linguistics.
It contains 17.7 million documents (5.1 billion tokens) compiled from 19 distinct archives, spanning 1,300 years from the 7th century to 2025. The corpus is linguistically diverse, covering Korean (Middle, Early Modern, Modern, North), Classical Chinese, and… See the full description on the dataset page: https://huggingface.co/datasets/seyoungsong/Open-Korean-Historical-Corpus.us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites
Historical US layoffs archive: 6,799 WARN Act notices that state websites no longer list (2000-2025), recovered
Rebuilt 2026-09-25. Five state labor agencies — Connecticut, Michigan, New York,
North Carolina and Pennsylvania — retired the web pages their older WARN Act
layoff notices lived on. Their current pages start years later. This dataset is
every notice in our file that came from one of those retired pages and is not
on the agency's live page today: 6,799 notices, 6,799… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites.wenhuibao_disk
文汇报光盘1938-1999
光盘收录了1938年-1999年所有文章共计1231692篇。13张光盘中包括扫描的图像和文本数据, setup.iso为安装程序(内含文本数据库),1-12文件夹为原来1-12号光碟。html.7z为爬虫爬取的文章html页面的压缩包。
光盘使用说明
安装程序需要在windows98 简体中文版中打开
进入系统后,插入setup.iso,安装时根据提示插入1-5号光盘;其他光盘在读取插图时使用。
(建议使用VirtualBox虚拟机)
启动爬虫
安装ie6
在安装目录的html/cgi-bin/oneart.htm的""后插入
<SCRIPT language=JavaScript>
var a =function() {
if (document.readyState !="complete") return;
var x = new ActiveXObject("Microsoft.XMLHTTP");
var content =… See the full description on the dataset page: https://huggingface.co/datasets/banned-historical-archives/wenhuibao_disk.legal-judgments
裁判文书 1985-
deepstock-stock-historical-prices-dataset-processedhuaqiaoribao
