datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SCBench
SCBench
[Paper]
[Code]
[Project Page]
SCBench (SharedContextBench) is a comprehensive benchmark to evaluate efficient long-context methods in a KV cache-centric perspective, analyzing their performance across the full KV cache lifecycle (generation, compression, retrieval, and loading) in real-world scenarios where context memory (KV cache) is shared and reused across multiple requests.
🎯 Quick Start
Load Data
You can download and load the SCBench data… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/SCBench.fineweb-edu-micro
FineWeb-Edu Micro
This dataset is a subset of the FineWeb-Edu Sample-10BT, which contains passages that are at least 1000 tokens long, totalling about 1 Million tokens .
This dataset was primarily made to evaluate different RAG Chunking mechanisms in Chonkie
Taskbench
TaskBench: Benchmarking Large Language Models for Task Automation
Introduction
TaskBench is a benchmark for evaluating large language models (LLMs) on task automation. Task automation can be formulated into three critical stages: task decomposition, tool invocation, and parameter prediction. This complexity makes data collection and evaluation more challenging compared to common NLP tasks. To address this challenge, we propose a comprehensive evaluation framework… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/Taskbench.btcusdt-microbar-v2
BTCUSDT Microbar v2
Sub-candle microstructure data for Binance USD-M Futures BTCUSDT, collected continuously over six WebSocket streams. Successor to Torch-Trade/btcusdt-microbar.
A standard OHLCV candle compresses thousands of trades into 6 numbers. This dataset preserves the raw event-level data — every individual trade, every best bid/ask change, every depth snapshot — so the underlying microstructure features can be reconstructed at any timeframe.
Why v2?
In April… See the full description on the dataset page: https://huggingface.co/datasets/Mindbyte-89/btcusdt-microbar-v2.R1_Lite_open_and_close_microwave_oven
R1_Lite_open_and_close_microwave_oven
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: galaxea_r1_lite
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
restaurant
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
push
pull
pressbutton… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/R1_Lite_open_and_close_microwave_oven.bing_coronavirus_query_set
Dataset Card for BingCoronavirusQuerySet
Dataset Summary
Please note that you can specify the start and end date of the data. You can get start and end dates from here: https://github.com/microsoft/BingCoronavirusQuerySet/tree/master/data/2020
example:
load_dataset("bing_coronavirus_query_set", queries_by="state", start_date="2020-09-01", end_date="2020-09-30")
You can also load the data by country by using queries_by="country".
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/bing_coronavirus_query_set.btc15m-market-microstructure
BTC15M Market Microstructure
This dataset is a research collection for Polymarket's 15-minute Bitcoin
UP/DOWN markets. It aligns public Polymarket market and order-book observations
with point-in-time Bitcoin market context from Binance and reference-price data
from Chainlink-related collection pipelines.
The package is designed to answer questions such as:
How do YES and NO prices react when Bitcoin moves USD 10, 20, 50, 100, or
more above or below the market reference price?… See the full description on the dataset page: https://huggingface.co/datasets/Pltrr/btc15m-market-microstructure.alpha_bot_2_operate_the_microwave_oven
alpha_bot_2_operate_the_microwave_oven
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: alpha_bot_2
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
pullapart
pushtogether
turn
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/alpha_bot_2_operate_the_microwave_oven.polymarket-updown-microstructure
Format. Three tables are published as parquet (under parquet/) for
the Hub viewer and pandas/polars/datasets users — pick a table from the
config dropdown above. The
honest-backtest loader
reads this parquet/ directory directly via its parquet adapter
(adapters.parquet_pm.load_corpus) for one-command reproduction of the
paper's results — see Reproduce the headline result below.
Dataset card — Polymarket crypto up/down microstructure (5m/15m)
Six weeks of real order-book… See the full description on the dataset page: https://huggingface.co/datasets/kinzikdza/polymarket-updown-microstructure.benchpress-score-matrix
BenchPress Score Matrix
This dataset contains the public model-by-benchmark score matrix used by
BenchPress. The release includes the lossless audited JSON, benchmark cost
evidence, flat model and benchmark metadata, one row per observed score, and
the paper-canonical dense subset used in the BenchPress experiments.
The source repository is
microsoft/benchpress.
Canonical artifacts
data/llm_benchmark_data.json is the authoritative rich score-matrix artifact.
It… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/benchpress-score-matrix.agibot-sim-heat-the-food-in-the-microwaveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "a2d",
"total_episodes": 75,
"total_frames": 172716,
"total_tasks": 1,
"total_videos": 225,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30.0,
"splits": {
"train": "0:75"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bot-pi/agibot-sim-heat-the-food-in-the-microwave.msr-acc-tae25
Microsoft Research - Accurate Chemistry Collection: Total Atomization Energies
Description
The Microsoft Research Accurate Chemistry Collection (MSR-ACC) provides a collection of accurate coupled cluster labels for training machine learning functionals.
MSR-ACC/TAE25 comprising 73,040 total atomization energies at the CCSD(T)/CBS level obtained with the W1-F12 thermochemical protocol.
The dataset is constructed to exhaustively cover the chemical space of closed-shell… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/msr-acc-tae25.microscope_pipette_2026-09-02This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"joint_1.pos",
"joint_2.pos",
"joint_3.pos",
"joint_4.pos",
"joint_5.pos",
"joint_6.pos",
"precision.state"
],
"shape": [… See the full description on the dataset page: https://huggingface.co/datasets/AdamAxelrod/microscope_pipette_2026-09-02.microscope_pipette_2026-09-10This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"joint_1.pos",
"joint_2.pos",
"joint_3.pos",
"joint_4.pos",
"joint_5.pos",
"joint_6.pos",
"precision.state"
],
"shape": [… See the full description on the dataset page: https://huggingface.co/datasets/AdamAxelrod/microscope_pipette_2026-09-10.tspgpn-msa-microglia-fullopen_microwave_25_08_21_lerobotv2.1CoSAlign-Train
CoSAlign-Train: A Large-Scale Synthetic Training Dataset for Controllable Safety Alignment
Paper: Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements, published at ICLR 2025.
Purpose: Training dataset for controllable safety alignment (CoSA) of large language models (LLMs), facilitating fine-grained inference-time adaptation to diverse safety requirements.
Description: CoSAlign-Train is a large-scale, synthetic preference dataset designed for… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/CoSAlign-Train.close_microwave_25_08_21_lerobotv2.1libero-data-for-rhoSS316L_small_10_micronSWE-bench_Verified_micro
SWE-bench_Verified_Micro
A subsampled version of SWE-Bench_Verified dataset with 50 samples, sampled using the original unique repo distribution from the original dataset
caged-microdados-traduzidos
CAGED — Microdados Traduzidos (Brasil, mercado completo)
Microdados do CAGED (Cadastro Geral de Empregados e Desempregados,
Ministério do Trabalho e Emprego) com os códigos traduzidos pelos
dicionários oficiais do próprio MTE.
O CAGED é publicado inteiramente codificado: sexo é 1, grau de instrução é
1..11, ocupação é um código CBO, setor é um código CNAE. Ler os microdados
crus exige cruzar à mão dezenas de planilhas de layout espalhadas pelo FTP do
ministério, que mudam de… See the full description on the dataset page: https://huggingface.co/datasets/Gianpedro/caged-microdados-traduzidos.portal-rebot-pick-microcontroller
portal-rebot-pick-microcontroller
A LeRobot dataset of a 6-DOF arm picking a
microcontroller and placing it in a box. Episodes were collected by teleoperating a real
robot over the network with LiveKit Portal, using a
human-in-the-loop recording loop that aligns camera frames and joint state into clean,
synchronized trajectories.
Task: Pick the microcontroller and place it in the box
At a glance
Robot
Seeed reBot Arm B601-DM (seeed_b601_dm_follower)… See the full description on the dataset page: https://huggingface.co/datasets/binhpham/portal-rebot-pick-microcontroller.dtwin_g1_microwave_bowlmicroscope_pipetteThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"joint_1.pos",
"joint_2.pos",
"joint_3.pos",
"joint_4.pos",
"joint_5.pos",
"joint_6.pos",
"precision.state"
],
"shape": [… See the full description on the dataset page: https://huggingface.co/datasets/AdamAxelrod/microscope_pipette.sattupperware-microwave
tupperware-microwave
LeRobot v2.1 format dataset for robot manipulation.
Dataset Structure
Episodes: 2 episodes of robot manipulation
Total Frames: 6152 frames
Cameras: 3 camera views per episode
observation.images.base_camera_sensor_image_raw
observation.images.arm1_camera_sensor_image_raw
observation.images.arm2_camera_sensor_image_raw
Robot: Bimanual manipulator with 34 joints
Format: LeRobot v2.1
Usage with LeRobot
from… See the full description on the dataset page: https://huggingface.co/datasets/Sraghvi/tupperware-microwave.openarm-bimanual-open-microwave-sim-iter-2-v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "OpenArm-Bimanual",
"total_episodes": 50,
"total_frames": 13561,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/crislmfroes/openarm-bimanual-open-microwave-sim-iter-2-v2.bio-faiss-microbiome-v1
bio-faiss-microbiome-v1
A FAISS index + metadata for scientific retrieval
Contents
index.faiss: FAISS index (cosine w/ inner product).
meta.jsonl: one JSON per chunk; fields include chunk_id, paper_id, title, section, subsection, paragraph_index, keywords, boost.
Build provenance
Chunking: hierarchical (section→paragraph→~380-token chunks, ~15% overlap)
Embedder: bio-protocol/scientific-retriever (mean-pooled, L2-normalized)
Similarity: cosine via inner… See the full description on the dataset page: https://huggingface.co/datasets/bio-protocol/bio-faiss-microbiome-v1.
