datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ethereum-arbitrage
Crypto & DeFi Documentation Dataset
A comprehensive dataset of cryptocurrency, DeFi, and blockchain documentation and code suitable for LLM training.
Dataset Description
This dataset contains scraped and processed documentation from various crypto/DeFi sources including:
Rust Ethereum libraries (ethers-rs, etc.)
Solidity documentation (official Solidity language docs)
Smart contracts (Uniswap, Aave, Balancer, SushiSwap, etc.)
Trading bots (MEV, flashloans, arbitrage)… See the full description on the dataset page: https://huggingface.co/datasets/Jcrandall541/ethereum-arbitrage.modality-conflict-arbitration-v2
Modality-Conflict Arbitration Benchmark (v2)
A controlled benchmark for studying how a vision-language model arbitrates between
its two input channels when they disagree — and whether that choice tracks the
reliability of each channel.
Each row is a single conflict trial: an image of one math problem paired with the
text of a different problem. Because the two ground-truth answers are carried side by
side, the model's output alone tells you which modality it followed — no… See the full description on the dataset page: https://huggingface.co/datasets/vlm-modality-research/modality-conflict-arbitration-v2.arb-instruct-v1
Arb-Agent Instruct v1
A specialized financial reasoning dataset containing 4,500+ Chain-of-Thought (CoT) Q&A pairs generated from SEC 10-K filings.
Unlike generic financial datasets that focus on simple extraction ("What was 2023 revenue?"), this dataset focuses on multi-hop reasoning, casual analysis, and risk assessment.
Dataset Statistics
Total Rows: 4,632 Source Documents: 100+ SEC 10-K Filings (2022-2024). Coverage: Top 50 S&P 500 companies across 8 sectors (Tech… See the full description on the dataset page: https://huggingface.co/datasets/ckerf/arb-instruct-v1.
