CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01futurehouse /lab-bench LAB-Bench The Language Agent Biology Benchmark, or LAB-Bench, is an evaluation dataset for AI systems intended to benchmark capabilities foundational to scientific research in biology. The dataset currently consists of 8 broad categories, comprising 30 narrower subtasks, including extracting information from the scientific literature (LitQA2), retrieving information from databases (DbQA) and supplementary information (SuppQA), reasoning about scientific figures (FigQA) and tables… See the full description on the dataset page: https://huggingface.co/datasets/futurehouse/lab-bench.imagequestion-answering1K<n<10K51 likes34k downloads1y agoHugging Face02linxy /USDT-M_Perpetual_Futures USDT-M Perpetual Futures (Binance) Binance USDT-margined perpetual futures historical data, including OHLCV klines, mark / index / premium-index prices, open-interest & long/short ratios, and funding rates. Auto-updated daily from the official Binance public data mirror. 币安 U 本位永续合约 历史数据集,包含 K 线、标记价格、指数价格、溢价指数、持仓量及多空比、资金费率, 每日从 Binance 官方数据镜像自动更新。 Last updated on 2026-09-23 09:03:37 UTC Usage Data is stored as Parquet in per-symbol subdirectories. import pandas as… See the full description on the dataset page: https://huggingface.co/datasets/linxy/USDT-M_Perpetual_Futures.texttime-series-forecasting1B<n<10B4 likes22k downloads1d agoHugging Face03futurehouse /BixBench BixBench Dataset Contains the dataset file BixBench.jsonl and each corresponding data capsule as a .zip file. Capsules are named CapsuleFolder-{uuid}.zip IMPORTANT UPDATE 2025/09/23: In ongoing work, we found that many questions in the original dataset had insufficient detail to be answerable, especially in the preferred open-answer setting. To address this, we've extensively re-reviewed and revised a substantial portion of the benchmark. We have also updated the format of the… See the full description on the dataset page: https://huggingface.co/datasets/futurehouse/BixBench.textn<1K41 likes14k downloads1y agoHugging Face04liaolw /3D-FUTUREtext1K<n<10K1 likes6.3k downloads9mo agoHugging Face05futuremoon /x_dataset_39 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/futuremoon/x_dataset_39.texttext-classification1B<n<10B2 likes6k downloads1y agoHugging Face06futurex-ai /Futurex-Past FutureX-Past 📜 Overview This repository contains a dataset of past questions from the FutureX benchmark. FutureX is a live, dynamic benchmark designed to evaluate the future prediction capabilities of Large Language Model (LLM) agents. It features a fully automated pipeline that generates new questions about upcoming real-world events, deploys agents to predict their outcomes, and scores the results automatically. For more information on the live benchmark… See the full description on the dataset page: https://huggingface.co/datasets/futurex-ai/Futurex-Past.textquestion-answering1K<n<10K8 likes3.6k downloads3d agoHugging Face07futurex-ai /Futurex-Online Submission Guidelines — Weekly Prediction Challenge 📄 Technical Report: https://arxiv.org/pdf/2508.11987 🌐 LeaderBoard: https://futurex-ai.github.io/ We run a weekly real-time prediction challenge. This repo always contains latest events. 1. Weekly Rules A new set of tasks is released every week. This week's tasks cover events with an end time between 2026-09-23, 24:00 (UTC+8) and 2026-09-29, 24:00 (UTC+8). You must download the tasks, make predictions, and… See the full description on the dataset page: https://huggingface.co/datasets/futurex-ai/Futurex-Online.textquestion-answeringn<1K26 likes1.8k downloads3d agoHugging Face08CoralLeiCN /hle_futurehouse_bronzeimagen<1K0 likes1.4k downloads1y agoHugging Face09predict-quant /binance-future-orderbooktabulartime-series-forecasting10M<n<100M1 likes1.3k downloads6mo agoHugging Face10OpenMOSS-Team /FutureOmni FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Predicting the future requires listening as well as seeing. 📖 Dataset Summary Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio–visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective understanding. FutureOmni is the first benchmark designed… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/FutureOmni.tabularquestion-answering1K<n<10K7 likes983 downloads8mo agoHugging Face11futurehouse /ether0-benchmark ether0-benchmark QA benchmark (test set) for the ether0 reasoning language model: https://huggingface.co/futurehouse/ether0 This benchmark is made from commonly used tasks - like reaction prediction in USPTO/ORD, molecular captioning from PubChem, or predicting GHS classification. It's unique from other benchmarks in that all answers are a molecule. It's balanced so that each task is about 25 questions, a reasonable amount for frontier model evaluations. The tasks generally follow… See the full description on the dataset page: https://huggingface.co/datasets/futurehouse/ether0-benchmark.textquestion-answeringn<1K17 likes779 downloads1y agoHugging Face12q-future /Q-Bench-HFimage1K<n<10K3 likes735 downloads2y agoHugging Face13Skywayne /Futures_202306_202312Future dada on FG and sc. tabular1M<n<10M0 likes728 downloads3y agoHugging Face14tradecatlabs /binance-futures-ohlcv-2018-2026 🐱 币安期货 Main4 数据集 (BTC/ETH/BNB/SOL) 本项目由交易猫基金会支持(交易猫基金会 CA:0x8a99b8d53eff6bc331af529af74ad267f3167777)。 精简版币安期货历史数据 - 只包含 4 个主要币种,适合快速下载和研究使用。 📊 数据概览 文件 记录数 压缩大小 时间范围 说明 candles_1m_main4_*.bin.zst 998 万 366 MB 2020-01 ~ 2026-01 1分钟K线 (Binary) futures_metrics_main4_*.bin.zst 152 万 49 MB 2021-12 ~ 2026-01 期货指标 (Binary) schema_*.sql.zst - 6.3 KB - TimescaleDB Schema 总计: 1150 万条记录,压缩后约 415MB 🎯 包含币种 币种 K线记录数 K线时间范围… See the full description on the dataset page: https://huggingface.co/datasets/tradecatlabs/binance-futures-ohlcv-2018-2026.tabulartime-series-forecasting100M<n<1B130 likes659 downloads7mo agoHugging Face15EdisonFU2025 /Domestic_Futures_tickstabular100M<n<1B2 likes550 downloads1y agoHugging Face16q-future /Q-Bench2-HFimage1K<n<10K1 likes519 downloads2y agoHugging Face17futurehouse /aviary-paper-datagatedTest split data uploaded in December 2024 for FutureHouse's aviary/ldp paper. textn<1K3 likes447 downloads2y agoHugging Face18danielharkin21 /futuressstabular100M<n<1B0 likes396 downloads5mo agoHugging Face19evalflow /hle_futurehouse_bronzeimagen<1K0 likes370 downloads1y agoHugging Face20delmiron27 /recorder-binance-futures-btc chronos recorder archive — Binance USDT-M futures BTC L2 order-book (depth snapshots + diffs) and trades streams for BTCUSDT and BTCUSDC perpetuals. Recorded 24/7 by the chronos market recorder (GitHub: BlackDigitalStudio/crypto-market-recorder) on VM scalper-recorder (Tokyo) up to 2026-08-05 ~18:xx UTC. Hourly snappy-parquet files, layout preserved 1:1 from gs://recorder-data-asia-0998ac51/chronos/scalper-recorder/... (GCP->HF migration 2026-08-22; ledger:… See the full description on the dataset page: https://huggingface.co/datasets/delmiron27/recorder-binance-futures-btc.tabular1B<n<10B0 likes352 downloads1mo agoHugging Face21lynx1231 /historical-futures-data-sample Historical Futures Data Sample This repository contains a free evaluation sample of historical futures data across selected contracts and frequencies. The complete catalog covers more than 2,000 futures roots and 900 million observations. View Data and Pricing: https://futuresforexandsomeindexes.com/ This package is a normalized evaluation sample containing 40 selected contracts across 8 futures roots: CL, ES, GC, SB, SR3, VX, ZC, ZN. The original root and contract files are… See the full description on the dataset page: https://huggingface.co/datasets/lynx1231/historical-futures-data-sample.tabular1M<n<10M1 likes350 downloads2mo agoHugging Face22FutureMa /EvasionBench EvasionBench EvasionBench is a benchmark dataset for detecting evasive answers in earnings call Q&A sessions. The task is to classify how directly corporate management addresses questions from financial analysts. Dataset Summary This dataset contains 16,726 question-answer pairs from earnings call transcripts, each labeled with one of three evasion levels. The labels were generated using the Eva-4B-V2 model, a fine-tuned classifier specifically… See the full description on the dataset page: https://huggingface.co/datasets/FutureMa/EvasionBench.texttext-classification10K<n<100K110 likes342 downloads7mo agoHugging Face23FutureMa /DramaBench DramaBench: Drama Script Continuation Dataset Dataset Summary DramaBench is a comprehensive benchmark dataset for evaluating drama script continuation capabilities of large language models. Current Release: v3.0 Full (1,103 samples) - The complete DramaBench collection is now openly available, with context-continuation pairs designed to assess models across six independent evaluation dimensions. Release Roadmap Version Samples Status… See the full description on the dataset page: https://huggingface.co/datasets/FutureMa/DramaBench.texttext-generation1K<n<10K111 likes332 downloads2mo agoHugging Face24futurehouse /hle-gold-bio-chemgated Humanity's Last Exam (HLE) Bio/Chem Gold Humanity’s Last Exam (HLE) is a challenging question-answering AI benchmark covering advanced academic fields including Math, Physics, Chemistry, Biology, Engineering, and Computer Science. At FutureHouse, we audited the biology and chemistry subsets of HLE using a combination of expert human evaluators and our in-house research agent, and found that around 30% of the questions contain answers directly contradicted by peer-reviewed… See the full description on the dataset page: https://huggingface.co/datasets/futurehouse/hle-gold-bio-chem.imagequestion-answeringn<1K25 likes291 downloads1y agoHugging Face25abdoelsayed /FutureQueryEval FutureQueryEval Dataset (EMNLP 2025)🔍 Dataset Description FutureQueryEval is a novel Information Retrieval (IR) benchmark designed to evaluate reranker performance on temporal novelty. It comprises 148 queries with 2,938 query-document pairs across 7 topical categories, specifically created to test how well reranking models generalize to truly novel queries that were unseen during LLM pretraining. Key Features Zero Contamination: All queries refer to events… See the full description on the dataset page: https://huggingface.co/datasets/abdoelsayed/FutureQueryEval.texttext-retrieval1K<n<10K2 likes287 downloads1y agoHugging Face26futureDoctor /turkic_tts_dataset Turkic TTS Dataset A multilingual TTS corpus covering Turkic languages. Languages Subset Source Speakers azerbaijani BHOSAI/Azerbaijani_News_TTS 1 (female) bashkir AigizK/bashkort_tts_dataset 8 (7F + 1M, ElevenLabs cloned) Schema Column Type Description audio Audio Speech sample text string Transcription source_link string Original dataset URL speaker_idstring Speaker identifier (and style if applicable) gender string… See the full description on the dataset page: https://huggingface.co/datasets/futureDoctor/turkic_tts_dataset.audiotext-to-speech10K<n<100K0 likes280 downloads4mo agoHugging Face27msj-21 /es-futures-1mtabular1M<n<10M0 likes238 downloads4mo agoHugging Face28evalflow /hle_futurehouse_goldimagen<1K0 likes224 downloads1y agoHugging Face29Khanhpham1992 /es-futures-1mtabular1M<n<10M0 likes223 downloads5mo agoHugging Face30future-architect /Llama-3.3-Future-Code-Instructions Llama 3.3 Future Code Instructions Llama 3.3 Future Code Instructions is a large-scale instruction dataset synthesized with the Meta Llama 3.3 70B Instruct model. The dataset was generated with the method called Magpie, where we prompted the model to generate instructions likely to be asked by the users. In addition to the original prompt introduced by the authors, we conditioned the system prompt on what specific programming language the user has an interest in, gaining control… See the full description on the dataset page: https://huggingface.co/datasets/future-architect/Llama-3.3-Future-Code-Instructions.text1M<n<10M0 likes160 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.