datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lab-bench
LAB-Bench
The Language Agent Biology Benchmark, or LAB-Bench, is an evaluation dataset for AI systems intended to benchmark capabilities foundational to scientific research in biology. The dataset currently consists of 8 broad categories, comprising 30 narrower subtasks, including extracting information from the scientific literature (LitQA2), retrieving information from databases (DbQA) and supplementary information (SuppQA), reasoning about scientific figures (FigQA) and tables… See the full description on the dataset page: https://huggingface.co/datasets/futurehouse/lab-bench.USDT-M_Perpetual_Futures
USDT-M Perpetual Futures (Binance)
Binance USDT-margined perpetual futures historical data, including OHLCV klines,
mark / index / premium-index prices, open-interest & long/short ratios, and funding rates.
Auto-updated daily from the official Binance public data mirror.
币安 U 本位永续合约 历史数据集,包含 K 线、标记价格、指数价格、溢价指数、持仓量及多空比、资金费率,
每日从 Binance 官方数据镜像自动更新。
Last updated on 2026-09-23 09:03:37 UTC
Usage
Data is stored as Parquet in per-symbol subdirectories.
import pandas as… See the full description on the dataset page: https://huggingface.co/datasets/linxy/USDT-M_Perpetual_Futures.BixBench
BixBench Dataset
Contains the dataset file BixBench.jsonl and each corresponding data capsule as a .zip file. Capsules are named CapsuleFolder-{uuid}.zip
IMPORTANT UPDATE 2025/09/23:
In ongoing work, we found that many questions in the original dataset had insufficient detail to be answerable, especially in the preferred open-answer setting. To address this, we've extensively re-reviewed and revised a substantial portion of the benchmark. We have also updated the format of the… See the full description on the dataset page: https://huggingface.co/datasets/futurehouse/BixBench.3D-FUTUREx_dataset_39
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/futuremoon/x_dataset_39.Futurex-Past
FutureX-Past
📜 Overview
This repository contains a dataset of past questions from the FutureX benchmark.
FutureX is a live, dynamic benchmark designed to evaluate the future prediction capabilities of Large Language Model (LLM) agents. It features a fully automated pipeline that generates new questions about upcoming real-world events, deploys agents to predict their outcomes, and scores the results automatically. For more information on the live benchmark… See the full description on the dataset page: https://huggingface.co/datasets/futurex-ai/Futurex-Past.Futurex-Online
Submission Guidelines — Weekly Prediction Challenge
📄 Technical Report: https://arxiv.org/pdf/2508.11987
🌐 LeaderBoard: https://futurex-ai.github.io/
We run a weekly real-time prediction challenge. This repo always contains latest events.
1. Weekly Rules
A new set of tasks is released every week.
This week's tasks cover events with an end time between 2026-09-23, 24:00 (UTC+8) and 2026-09-29, 24:00 (UTC+8).
You must download the tasks, make predictions, and… See the full description on the dataset page: https://huggingface.co/datasets/futurex-ai/Futurex-Online.hle_futurehouse_bronzebinance-future-orderbookFutureOmni
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
Predicting the future requires listening as well as seeing.
📖 Dataset Summary
Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio–visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective understanding.
FutureOmni is the first benchmark designed… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/FutureOmni.ether0-benchmark
ether0-benchmark
QA benchmark (test set) for the ether0 reasoning language model:
https://huggingface.co/futurehouse/ether0
This benchmark is made from commonly used tasks - like reaction prediction in USPTO/ORD,
molecular captioning from PubChem, or predicting GHS classification.
It's unique from other benchmarks in that all answers are a molecule.
It's balanced so that each task is about 25 questions,
a reasonable amount for frontier model evaluations.
The tasks generally follow… See the full description on the dataset page: https://huggingface.co/datasets/futurehouse/ether0-benchmark.Q-Bench-HFFutures_202306_202312Future dada on FG and sc.
binance-futures-ohlcv-2018-2026
🐱 币安期货 Main4 数据集 (BTC/ETH/BNB/SOL)
本项目由交易猫基金会支持(交易猫基金会 CA:0x8a99b8d53eff6bc331af529af74ad267f3167777)。
精简版币安期货历史数据 - 只包含 4 个主要币种,适合快速下载和研究使用。
📊 数据概览
文件
记录数
压缩大小
时间范围
说明
candles_1m_main4_*.bin.zst
998 万
366 MB
2020-01 ~ 2026-01
1分钟K线 (Binary)
futures_metrics_main4_*.bin.zst
152 万
49 MB
2021-12 ~ 2026-01
期货指标 (Binary)
schema_*.sql.zst
-
6.3 KB
-
TimescaleDB Schema
总计: 1150 万条记录,压缩后约 415MB
🎯 包含币种
币种
K线记录数
K线时间范围… See the full description on the dataset page: https://huggingface.co/datasets/tradecatlabs/binance-futures-ohlcv-2018-2026.Domestic_Futures_ticksQ-Bench2-HFaviary-paper-dataTest split data uploaded in December 2024
for FutureHouse's aviary/ldp paper.
futuressshle_futurehouse_bronzerecorder-binance-futures-btc
chronos recorder archive — Binance USDT-M futures BTC
L2 order-book (depth snapshots + diffs) and trades streams for BTCUSDT and BTCUSDC perpetuals.
Recorded 24/7 by the chronos market recorder (GitHub:
BlackDigitalStudio/crypto-market-recorder) on VM scalper-recorder (Tokyo) up to
2026-08-05 ~18:xx UTC. Hourly snappy-parquet files, layout preserved 1:1 from
gs://recorder-data-asia-0998ac51/chronos/scalper-recorder/...
(GCP->HF migration 2026-08-22; ledger:… See the full description on the dataset page: https://huggingface.co/datasets/delmiron27/recorder-binance-futures-btc.historical-futures-data-sample
Historical Futures Data Sample
This repository contains a free evaluation sample of historical futures data across selected contracts and frequencies.
The complete catalog covers more than 2,000 futures roots and 900 million observations.
View Data and Pricing: https://futuresforexandsomeindexes.com/
This package is a normalized evaluation sample containing 40 selected contracts across 8 futures roots: CL, ES, GC, SB, SR3, VX, ZC, ZN.
The original root and contract files are… See the full description on the dataset page: https://huggingface.co/datasets/lynx1231/historical-futures-data-sample.EvasionBench
EvasionBench
EvasionBench is a benchmark dataset for detecting evasive answers in earnings call Q&A sessions. The task is to classify how directly corporate management addresses questions from financial analysts.
Dataset Summary
This dataset contains 16,726 question-answer pairs from earnings call transcripts, each labeled with one of three evasion levels. The labels were generated using the Eva-4B-V2 model, a fine-tuned classifier specifically… See the full description on the dataset page: https://huggingface.co/datasets/FutureMa/EvasionBench.DramaBench
DramaBench: Drama Script Continuation Dataset
Dataset Summary
DramaBench is a comprehensive benchmark dataset for evaluating drama script continuation capabilities of large language models.
Current Release: v3.0 Full (1,103 samples) - The complete DramaBench collection is now openly available, with context-continuation pairs designed to assess models across six independent evaluation dimensions.
Release Roadmap
Version
Samples
Status… See the full description on the dataset page: https://huggingface.co/datasets/FutureMa/DramaBench.hle-gold-bio-chem
Humanity's Last Exam (HLE) Bio/Chem Gold
Humanity’s Last Exam (HLE) is a challenging question-answering AI benchmark covering advanced academic fields including Math, Physics, Chemistry, Biology, Engineering, and Computer Science.
At FutureHouse, we audited the biology and chemistry subsets of HLE using a combination of expert human evaluators and our in-house research agent, and found that around 30% of the questions contain answers directly contradicted by peer-reviewed… See the full description on the dataset page: https://huggingface.co/datasets/futurehouse/hle-gold-bio-chem.FutureQueryEval
FutureQueryEval Dataset (EMNLP 2025)🔍
Dataset Description
FutureQueryEval is a novel Information Retrieval (IR) benchmark designed to evaluate reranker performance on temporal novelty. It comprises 148 queries with 2,938 query-document pairs across 7 topical categories, specifically created to test how well reranking models generalize to truly novel queries that were unseen during LLM pretraining.
Key Features
Zero Contamination: All queries refer to events… See the full description on the dataset page: https://huggingface.co/datasets/abdoelsayed/FutureQueryEval.turkic_tts_dataset
Turkic TTS Dataset
A multilingual TTS corpus covering Turkic languages.
Languages
Subset
Source
Speakers
azerbaijani
BHOSAI/Azerbaijani_News_TTS
1 (female)
bashkir
AigizK/bashkort_tts_dataset
8 (7F + 1M, ElevenLabs cloned)
Schema
Column
Type
Description
audio
Audio
Speech sample
text
string
Transcription
source_link
string
Original dataset URL
speaker_idstring
Speaker identifier (and style if applicable)
gender
string… See the full description on the dataset page: https://huggingface.co/datasets/futureDoctor/turkic_tts_dataset.es-futures-1mhle_futurehouse_goldes-futures-1mLlama-3.3-Future-Code-Instructions
Llama 3.3 Future Code Instructions
Llama 3.3 Future Code Instructions is a large-scale instruction dataset synthesized with the Meta Llama 3.3 70B Instruct model.
The dataset was generated with the method called Magpie, where we prompted the model to generate instructions likely to be asked by the users.
In addition to the original prompt introduced by the authors, we conditioned the system prompt on what specific programming language the user has an interest in, gaining control… See the full description on the dataset page: https://huggingface.co/datasets/future-architect/Llama-3.3-Future-Code-Instructions.
