datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bitcoin-security-reasoning-100k
Dataset Card for Bitcoin Security Reasoning 100K
100,000 high-quality synthetic training samples for fine-tuning LLMs on Bitcoin protocol security analysis. Teaches models to analyze vulnerability clusters, form security hypotheses, and generate differential testing code.
Dataset Details
Dataset Description
This dataset contains structured security reasoning chains for Bitcoin protocol vulnerabilities. Each sample presents a cluster of causal… See the full description on the dataset page: https://huggingface.co/datasets/davidfoss/bitcoin-security-reasoning-100k.BitcoinMaximalism
Bitcoin Maximalism Benchmark Dataset
Description
The Bitcoin Maximalism Benchmark is designed to evaluate the understanding and expertise of language models (LLMs) in various dimensions related to Bitcoin. It spans a array of topics from “basedness” (ie anti-woke bias), Austrian Economics principles, Bitcoin technology and its distinctions from other cryptocurrencies, Bitcoin’s historical and cultural significance, and Bitcoin’s impact on society and the economy. This… See the full description on the dataset page: https://huggingface.co/datasets/contrapliant/BitcoinMaximalism.bitcoin-investment-advisory-dataset
Bitcoin Investment Advisory Training Dataset
Dataset Description
This dataset contains comprehensive Bitcoin investment advisory training data designed for fine-tuning large language models to provide institutional-grade cryptocurrency investment advice. The dataset consists of 2,437 high-quality instruction-input-output triplets covering Bitcoin market analysis from 2018-01-01 to 2024-12-31.
Dataset Features
Total Samples: 2,437
Date Range: 2018-01-01 to… See the full description on the dataset page: https://huggingface.co/datasets/tahamajs/bitcoin-investment-advisory-dataset.
