datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Blockchain-Sensitive-Detect-Data
Blockchain-Sensitive-Detect-Data
English README
复旦大学附属儿科医院-区块链敏感信息检测项目的多模态完整测试数据集。
项目仓库:https://github.com/anyangsong/Blockchain-Sensitive-Detect
数据集:https://huggingface.co/datasets/anyangsong/Blockchain-Sensitive-Detect-Data
checkpoints:https://huggingface.co/anyangsong/Blockchain-Sensitive-Detect-Checkpoints
数据以原始文件夹组织,覆盖文本、音频、图像与视频等样本。
该仓库不提供统一的 CSV、Parquet 或 JSONL 清单;类别信息主要由目录名和文件名携带。
内容警告: 数据集包含辱骂、性内容、暴力、政治相关内容、误导性医疗信息、欺诈信息。使用者应仅在具备适当访问控制、伦理审查和当地法律依据的环境中处理这些内容。… See the full description on the dataset page: https://huggingface.co/datasets/anyangsong/Blockchain-Sensitive-Detect-Data.lm-eval-results-alnrg2arg-blockchainlabs_7B_merged_test2_4-private
Dataset Card for Evaluation run of alnrg2arg/blockchainlabs_7B_merged_test2_4
Dataset automatically created during the evaluation run of model alnrg2arg/blockchainlabs_7B_merged_test2_4
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-alnrg2arg-blockchainlabs_7B_merged_test2_4-private.blockchain-benchmark
Dataset Card for LLM Blockchain Benchmark
Dataset Summary
The Blockchain Benchmark Dataset is a comprehensive collection of data specifically curated for benchmarking Language Models (LMs) in the domain of blockchain technology. This dataset is designed to facilitate research and development in natural language understanding within the blockchain domain.
A complete list of tasks: ['general-reasoning', 'code', 'math']
Supported Tasks and Leaderboards
Model… See the full description on the dataset page: https://huggingface.co/datasets/revflask/blockchain-benchmark.lm-eval-results-alnrg2arg-blockchainlabs_test3_seminar-private
Dataset Card for Evaluation run of alnrg2arg/blockchainlabs_test3_seminar
Dataset automatically created during the evaluation run of model alnrg2arg/blockchainlabs_test3_seminar
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-alnrg2arg-blockchainlabs_test3_seminar-private.Ethereum_blockchain_parquet
Data source
The blockchain data was extracted into parquet files using the Rust-based
cryo tool in combination with the web3 API
provided by Ankr. Note that an account needs to be
created even if you intend to make use of the limited capabilities under the freemium
pricing model.
Sample command for using cryo is shown below. It can take long (~ 2 hours) to complete
since we need to stay below the rate limit of 30 reqs/min. Also, we are fetching every
25th block, which means the data… See the full description on the dataset page: https://huggingface.co/datasets/vnegi10/Ethereum_blockchain_parquet.details_alnrg2arg__blockchainlabs_tinyllama_fusion_LHK_yunkong
Dataset Card for Evaluation run of alnrg2arg/blockchainlabs_tinyllama_fusion_LHK_yunkong
Dataset automatically created during the evaluation run of model alnrg2arg/blockchainlabs_tinyllama_fusion_LHK_yunkong on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_alnrg2arg__blockchainlabs_tinyllama_fusion_LHK_yunkong.details_alnrg2arg__blockchainlabs_tinyllama_fusion_LHK_yunkong_v2
Dataset Card for Evaluation run of alnrg2arg/blockchainlabs_tinyllama_fusion_LHK_yunkong_v2
Dataset automatically created during the evaluation run of model alnrg2arg/blockchainlabs_tinyllama_fusion_LHK_yunkong_v2 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_alnrg2arg__blockchainlabs_tinyllama_fusion_LHK_yunkong_v2.omie-blockchain-study
OMIE Intraday Auction Subsets for Blockchain Governance Study
This repository contains cleaned and size-scaled tabular datasets derived from one representative day of the OMIE European intraday auction market. The data were prepared for experiments on blockchain-based verifiable governance of electricity flexibility and auction markets.
The dataset is designed for three uses:
replaying a realistic intraday auction day end-to-end,
benchmarking a blockchain-governed architecture… See the full description on the dataset page: https://huggingface.co/datasets/UniversitatdeLleida/omie-blockchain-study.cyberattack-blockchain-synth
ELTEX-Blockchain: A Domain-Specific Dataset for Cybersecurity
🔐 12k Synthetic Social Media Messages for Early Cyberattack Detection on Blockchain
Dataset Statistics
Category
Samples
Description
Cyberattack
6,941
Early warning signals and indicators of cyberattacks
General
4,507
Regular blockchain discussions (non-security related)
Dataset Structure
Each entry in the dataset contains:
message_id: Unique identifier for each message… See the full description on the dataset page: https://huggingface.co/datasets/dn-institute/cyberattack-blockchain-synth.details_alnrg2arg__blockchainlabs_7B_merged_test2_4
Dataset Card for Evaluation run of alnrg2arg/blockchainlabs_7B_merged_test2_4
Dataset automatically created during the evaluation run of model alnrg2arg/blockchainlabs_7B_merged_test2_4 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_alnrg2arg__blockchainlabs_7B_merged_test2_4.blockchain-sim-2sapientblock-blockchain-use-cases
SapientBlock Blockchain Use Cases
Der Datensatz enthält 255 redaktionell geprüfte Blockchain-Use-Cases aus 74 Branchen. Er stellt die öffentlich zugänglichen SapientBlock-Inhalte in einem maschinenlesbaren JSONL-Format für Forschung, Bildung, Retrieval und Quellenanalyse bereit.
SapientBlock ist ein öffentliches Forschungs- und Bildungsprojekt von ShapeNeural. Die Inhalte sind keine Rechts-, Investitions-, Unternehmens- oder technische Beratung.
Inhalt
Jeder… See the full description on the dataset page: https://huggingface.co/datasets/TooKeen/sapientblock-blockchain-use-cases.Solana-blockchain-360-CodingThis dataset contains 360 coding and tech related samples for the Solana blockchain.
Language: English
Coding-languages: Rust, Typescript, & C#
214 general knowledge samples
146 coding knowledge samples
Dataset Catalog:
201 Solana blockchain knowledge samples
49 Solana typescript coding samples
86 Solana rust coding samples
13 Solnet SDK knowledge samples
11 Solana c# coding samples
details_alnrg2arg__blockchainlabs_joe_bez_seminar
Dataset Card for Evaluation run of alnrg2arg/blockchainlabs_joe_bez_seminar
Dataset automatically created during the evaluation run of model alnrg2arg/blockchainlabs_joe_bez_seminar on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_alnrg2arg__blockchainlabs_joe_bez_seminar.omie_blockchain_study
omie_blockchain_study (TsFile)
Apache TsFile version of UniversitatdeLleida/omie-blockchain-study.
Converted rows: 2,367
Data files: ['omie_blockchain_study.tsfile']
Usage
Install the Apache TsFile Python SDK (pip install tsfile) and read a converted file:
from pathlib import Path
from tsfile import TsFileReader
path = Path("omie_blockchain_study.tsfile")
with TsFileReader(str(path)) as reader:
schemas = reader.get_all_table_schemas()
print("tables:"… See the full description on the dataset page: https://huggingface.co/datasets/THULab/omie_blockchain_study.blockchain-simblockchain-benchmark-formatted
Dataset Card for LLM Blockchain Benchmark
Dataset Summary
The Blockchain Benchmark Dataset is a comprehensive collection of data specifically curated for benchmarking Language Models (LMs) in the domain of blockchain technology. This dataset is designed to facilitate research and development in natural language understanding within the blockchain domain.
A complete list of tasks: ['general-reasoning', 'code', 'math']
Supported Tasks and Leaderboards
Model… See the full description on the dataset page: https://huggingface.co/datasets/revflask/blockchain-benchmark-formatted.details_alnrg2arg__blockchainlabs_7B_merged_test2_4_prune
Dataset Card for Evaluation run of alnrg2arg/blockchainlabs_7B_merged_test2_4_prune
Dataset automatically created during the evaluation run of model alnrg2arg/blockchainlabs_7B_merged_test2_4_prune on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_alnrg2arg__blockchainlabs_7B_merged_test2_4_prune.details_alnrg2arg__blockchainlabs_test3_seminar
Dataset Card for Evaluation run of alnrg2arg/blockchainlabs_test3_seminar
Dataset automatically created during the evaluation run of model alnrg2arg/blockchainlabs_test3_seminar on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_alnrg2arg__blockchainlabs_test3_seminar.paperzilla-blockchain-250
Paperzilla Blockchain Research Benchmark (250 papers, 5 LLM annotators)
Dataset Description
A multi-annotator benchmark dataset for evaluating retrieval systems on blockchain and cryptocurrency research papers. This dataset contains 250 cryptography and security papers from arXiv (cs.CR category), each independently annotated by 5 different large language models for relevance to blockchain research from a scientific perspective.
Key Features
250 papers from… See the full description on the dataset page: https://huggingface.co/datasets/paperzilla/paperzilla-blockchain-250.Machine-Checkable-Blockchain-Execution-Specification
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BearNetworkChain/Machine-Checkable-Blockchain-Execution-Specification.Machine-Checkable-Blockchain-Execution-Specification
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BNES-BRNKC/Machine-Checkable-Blockchain-Execution-Specification.bankless_What_is_Celestia__TIA__Unpacking_Modular_Blockchainsblockchain-treeblockchain-smartcontracts-1000
Blockchain y Contratos Inteligentes Dataset 1000
Dataset de 1000 instrucciones sobre los fundamentos de Blockchain, Contratos Inteligentes, DAOs, DeFi y programación en Solidity.
Uso
from datasets import load_dataset
dataset = load_dataset("miguelmejias0512/solidity_personal_dataset")
---
## Licencia
CC-BY-4.0 - Uso educativo
blockchain-research-training
Blockchain Research Training Data
On-chain data research dataset spanning Bitcoin Ordinals, BRC-20 inscriptions, Arweave, and Solana content. Raw inscription metadata, content analysis, and training-ready text for blockchain AI research.
Dataset Summary
Metric
Value
Chains
Bitcoin, Arweave, Solana
Unified Records
135 (bitcoin: 127, arweave: 5, solana: 3)
Training Records
100 (text-based inscriptions)
Content Types
PNG, WebP, GIF, SVG, HTML, MP3, MP4, JSON… See the full description on the dataset page: https://huggingface.co/datasets/purplesquirrelnetworks/blockchain-research-training.blockchain-leads-the-complete-package
Blockchain Leads: The Complete Package — Market Intelligence Dataset
40,000 verified Worldwide contacts are available from LeadsBlue →. This open dataset provides the aggregate market intelligence behind that database — contact volume, benchmark open/reply rates, send timing, and compliance for the Worldwide segment.
At a glance: A research dataset describing the Blockchain Leads: The Complete Package market: verified-contact volume, industry distribution, outreach benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/emailmarketingdataset/blockchain-leads-the-complete-package.cyberattack-blockchain-synth
ELTEX-Blockchain: A Domain-Specific Dataset for Cybersecurity
🔐 12k Synthetic Social Media Messages for Early Cyberattack Detection on Blockchain
Dataset Statistics
Category
Samples
Description
Cyberattack
6,941
Early warning signals and indicators of cyberattacks
General
4,507
Regular blockchain discussions (non-security related)
Dataset Structure
Each entry in the dataset contains:
message_id: Unique identifier for each message… See the full description on the dataset page: https://huggingface.co/datasets/cmp81/cyberattack-blockchain-synth.nigerian_transport_and_logistics_blockchain_logistics
Nigeria Transport & Logistics – Blockchain Logistics Records | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: infrastructure_transport - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/nigerian_transport_and_logistics_blockchain_logistics.bankless_EARLY_ACCESS_The_Blockchain_Trilemma_-_ETH_Vs_SOL_Vs_ATOM_with_Mike_Ippolito
