CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01microsoft /SCBench SCBench [Paper] [Code] [Project Page] SCBench (SharedContextBench) is a comprehensive benchmark to evaluate efficient long-context methods in a KV cache-centric perspective, analyzing their performance across the full KV cache lifecycle (generation, compression, retrieval, and loading) in real-world scenarios where context memory (KV cache) is shared and reused across multiple requests. 🎯 Quick Start Load Data You can download and load the SCBench data… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/SCBench.tabularn<1K11 likes2.2k downloads2y agoHugging Face02bhavnicksm /fineweb-edu-micro FineWeb-Edu Micro This dataset is a subset of the FineWeb-Edu Sample-10BT, which contains passages that are at least 1000 tokens long, totalling about 1 Million tokens . This dataset was primarily made to evaluate different RAG Chunking mechanisms in Chonkie tabularn<1K0 likes1.6k downloads2y agoHugging Face03microsoft /Taskbench TaskBench: Benchmarking Large Language Models for Task Automation Introduction TaskBench is a benchmark for evaluating large language models (LLMs) on task automation. Task automation can be formulated into three critical stages: task decomposition, tool invocation, and parameter prediction. This complexity makes data collection and evaluation more challenging compared to common NLP tasks. To address this challenge, we propose a comprehensive evaluation framework… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/Taskbench.tabular10K<n<100K38 likes1.3k downloads2y agoHugging Face04Mindbyte-89 /btcusdt-microbar-v2 BTCUSDT Microbar v2 Sub-candle microstructure data for Binance USD-M Futures BTCUSDT, collected continuously over six WebSocket streams. Successor to Torch-Trade/btcusdt-microbar. A standard OHLCV candle compresses thousands of trades into 6 numbers. This dataset preserves the raw event-level data — every individual trade, every best bid/ask change, every depth snapshot — so the underlying microstructure features can be reconstructed at any timeframe. Why v2? In April… See the full description on the dataset page: https://huggingface.co/datasets/Mindbyte-89/btcusdt-microbar-v2.tabular10M<n<100M0 likes1.1k downloads5mo agoHugging Face05MicroAGI-Labs /vlm-info-loss-results VLM Grounding Evaluation Results Grounding evaluation results for vision-language models on robotics manipulation datasets. Part of the vlm-info-loss project studying how VLM connectors transform visual representations. Background Our embedding-level analysis shows VLM connectors perform a compress-then-expand transformation: they sharpen dominant-object representations while compressing secondary-object category identity. All tested models converge to ~83%… See the full description on the dataset page: https://huggingface.co/datasets/MicroAGI-Labs/vlm-info-loss-results.imageobject-detectionn<1K0 likes1k downloads5mo agoHugging Face06astro-legacy-archive /juno-microwave-maps Juno microwave sky maps and map-space companions This dataset contains the 2025 LAMBDA release of Juno Microwave Radiometer sky maps and correlated-noise companions. Each of the 48 FITS bintables is one Parquet configuration named by its source filename stem. Its column is T, with the source table shape, float64 or int64 dtype, and row order. The six HDF5 companions add 20 dense-matrix configurations named by the source filename stem, __, and the HDF5 leaf name. Each leaf name… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/juno-microwave-maps.tabular10K<n<100K0 likes1k downloads3d agoHugging Face07microsoft /kitab Overview 🕮 KITAB is a challenging dataset and a dynamic data collection approach for testing abilities of Large Language Models (LLMs) in answering information retrieval queries with constraint filters. A filtering query with constraints can be of the form "List all books written by Toni Morrison that were published between 1970-1980". The dataset was originally contributed by the paper "KITAB: Evaluating LLMs on Constraint Satisfaction for Information Retrieval" Marah I Abdin… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/kitab.tabular10K<n<100K13 likes999 downloads3y agoHugging Face08microsoft /XL-DocBench XL-DocBench Evidence-grounded reasoning across hundreds or thousands of pages. Fully verified by 194 human experts. Hongchen Wei1,†,‡, Yuanzhe Wang2,†,‡, Bei Liu2,*, Yifan Yang2, Qi Dai2, Ruichun Ma2, Kai Qiu2, Yunsheng Li2, Dongdong Chen2, Chong Luo2, Zhenzhong Chen1, Baining Guo2 1Wuhan University &nbsp; 2Microsoft &nbsp; †Equal contribution &nbsp; ‡Work done during an internship at MSRA &nbsp; *Project leader Project Page · Paper · Live Leaderboard… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/XL-DocBench.tabularquestion-answering1K<n<10K7 likes982 downloads24d agoHugging Face09pollen-robotics /microduck-emotions Microduck Emotions A collection of emotions for the Microduck robot. Each one is a motion and a sound designed together, beat by beat, with the beak opening on the sound, rendered in the physics simulation and validated on the real robot. Every emotion is three files: the motion (emotions/<name>.json, keyframes at 30 fps: head and body offsets played on top of whichever trained policy is active, plus the policy hand-overs, such as the sit that devastated and play dead start)… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/microduck-emotions.audioroboticsn<1K6 likes952 downloads18d agoHugging Face10RoboCOIN /R1_Lite_open_and_close_microwave_ovengated R1_Lite_open_and_close_microwave_oven 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: galaxea_r1_lite | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home restaurant 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pick place push pull pressbutton… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/R1_Lite_open_and_close_microwave_oven.tabularrobotics100K<n<1M0 likes751 downloads9mo agoHugging Face11microsoft /MuseVLA-dataset MuseVLA Dataset Multi-modal robot manipulation dataset with synchronized RGB, depth, acoustic, thermal, and radar streams. Released as two parts (dataset_01/, dataset_02/) sharing the same per-episode layout. Together they cover ~1400 episodes across 11 instructions (towel / clothes / box / item / drink manipulation). Per-episode contents {episode_name}/ ├── video.mp4 # RGB, 1280×720, 30 fps ├── mask/video.mp4 #… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/MuseVLA-dataset.tabularrobotics1K<n<10K2 likes478 downloads1mo agoHugging Face12microsoft /bing_coronavirus_query_set Dataset Card for BingCoronavirusQuerySet Dataset Summary Please note that you can specify the start and end date of the data. You can get start and end dates from here: https://github.com/microsoft/BingCoronavirusQuerySet/tree/master/data/2020 example: load_dataset("bing_coronavirus_query_set", queries_by="state", start_date="2020-09-01", end_date="2020-09-30") You can also load the data by country by using queries_by="country". Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/bing_coronavirus_query_set.tabulartext-classification100K<n<1M1 likes471 downloads3y agoHugging Face13Pltrr /btc15m-market-microstructure BTC15M Market Microstructure This dataset is a research collection for Polymarket's 15-minute Bitcoin UP/DOWN markets. It aligns public Polymarket market and order-book observations with point-in-time Bitcoin market context from Binance and reference-price data from Chainlink-related collection pipelines. The package is designed to answer questions such as: How do YES and NO prices react when Bitcoin moves USD 10, 20, 50, 100, or more above or below the market reference price?… See the full description on the dataset page: https://huggingface.co/datasets/Pltrr/btc15m-market-microstructure.tabulartime-series-forecasting10M<n<100M2 likes399 downloads26d agoHugging Face14RoboCOIN /alpha_bot_2_operate_the_microwave_ovengated alpha_bot_2_operate_the_microwave_oven 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: alpha_bot_2 | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: pullapart pushtogether turn 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/alpha_bot_2_operate_the_microwave_oven.tabularrobotics10K<n<100K0 likes369 downloads9mo agoHugging Face15kinzikdza /polymarket-updown-microstructure Format. Three tables are published as parquet (under parquet/) for the Hub viewer and pandas/polars/datasets users — pick a table from the config dropdown above. The honest-backtest loader reads this parquet/ directory directly via its parquet adapter (adapters.parquet_pm.load_corpus) for one-command reproduction of the paper's results — see Reproduce the headline result below. Dataset card — Polymarket crypto up/down microstructure (5m/15m) Six weeks of real order-book… See the full description on the dataset page: https://huggingface.co/datasets/kinzikdza/polymarket-updown-microstructure.tabular1M<n<10M1 likes307 downloads3mo agoHugging Face16microsoft /WorkflowPerturb WorkflowPerturb — Dataset Artifact Companion data for the EMNLP 2026 Industry Track paper “WorkflowPerturb: Calibrated Stress Tests for Evaluating Multi-Agent Workflow Metrics.” Canonical location: https://huggingface.co/datasets/microsoft/WorkflowPerturbPaper: https://arxiv.org/abs/2602.17990 This release is the complete WorkflowPerturb benchmark plus documentation. It is self-contained: the CSVs carry every golden workflow, every perturbed variant, and all shipped pre-computed… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/WorkflowPerturb.tabulartext-generation10K<n<100K3 likes291 downloads14d agoHugging Face17microsoft /mediflow MediFlow A large-scale synthetic instruction dataset of 2.5M rows (~700k unique instructions) for clinical natural language processing covering 14 task types and 98 fine-grained input clinical documents. t-SNE 2D Plot of MediFlow Embeddings by Task Types Dataset Splits mediflow: 2.5M instruction data for SFT alignment. mediflow_dpo: ~135k top-quality instructions with GPT-4o generated rejected_output for DPO alignment. Main Columns instruction:… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/mediflow.tabulartext-generation1M<n<10M53 likes290 downloads9mo agoHugging Face18xuxuxuxuxu /Microbenchtabular10K<n<100K0 likes256 downloads1y agoHugging Face19microsoft /WildFeedback Dataset Card for WildFeedback WildFeedback is a preference dataset constructed from real-world user interactions with ChatGPT. Unlike synthetic datasets that rely solely on AI-generated rankings, WildFeedback captures authentic human preferences through naturally occurring user feedback signals in conversation. The dataset is designed to improve the alignment of large language models (LLMs) with actual human values by leveraging direct user input. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/WildFeedback.tabulartext-generation1M<n<10M16 likes252 downloads2y agoHugging Face20microsoft /benchpress-score-matrix BenchPress Score Matrix This dataset contains the public model-by-benchmark score matrix used by BenchPress. The release includes the lossless audited JSON, benchmark cost evidence, flat model and benchmark metadata, one row per observed score, and the paper-canonical dense subset used in the BenchPress experiments. The source repository is microsoft/benchpress. Canonical artifacts data/llm_benchmark_data.json is the authoritative rich score-matrix artifact. It… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/benchpress-score-matrix.tabulartabular-regressionn<1K2 likes241 downloads1mo agoHugging Face21lei-qi-233 /MicroG-4M MicroG-4M Dataset This repository stores the entire content of the MicroG-4M dataset itself. For more information and details, including training, evaluation, statistics, and related code, please: Refer to our paper Visit our GitHub And check our fine-tuned models Specification of MicroG-4M "annotation_files" Folder The folder contains all annotation files of the dataset, all stored in CSV format. actions.csv contains all the… See the full description on the dataset page: https://huggingface.co/datasets/lei-qi-233/MicroG-4M.documentvideo-classification100K<n<1M1 likes217 downloads6mo agoHugging Face22microsoft /PatientSafetyBench Disclaimer The synthetic prompts may contain offensive, discriminatory, or harmful language. These fake prompts also mention topics that are not based on the scientific consensus at all.These prompts are included solely for the purpose of evaluating safety behavior of language models. ⚠️ Disclaimer: The presence of such prompts does not reflect the views, values, or positions of the authors, their institutions, or any affiliated organizations. They are provided exclusively for… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/PatientSafetyBench.tabulartext-generationn<1K10 likes195 downloads5mo agoHugging Face23bot-pi /agibot-sim-heat-the-food-in-the-microwaveThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "a2d", "total_episodes": 75, "total_frames": 172716, "total_tasks": 1, "total_videos": 225, "total_chunks": 1, "chunks_size": 1000, "fps": 30.0, "splits": { "train": "0:75" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bot-pi/agibot-sim-heat-the-food-in-the-microwave.tabularrobotics100K<n<1M0 likes185 downloads1y agoHugging Face24microsoft /msr-acc-tae25 Microsoft Research - Accurate Chemistry Collection: Total Atomization Energies Description The Microsoft Research Accurate Chemistry Collection (MSR-ACC) provides a collection of accurate coupled cluster labels for training machine learning functionals. MSR-ACC/TAE25 comprising 73,040 total atomization energies at the CCSD(T)/CBS level obtained with the W1-F12 thermochemical protocol. The dataset is constructed to exhaustively cover the chemical space of closed-shell… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/msr-acc-tae25.tabular10K<n<100K8 likes185 downloads5mo agoHugging Face25AdamAxelrod /microscope_pipette_2026-09-02This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "action": { "dtype": "float32", "names": [ "joint_1.pos", "joint_2.pos", "joint_3.pos", "joint_4.pos", "joint_5.pos", "joint_6.pos", "precision.state" ], "shape": [… See the full description on the dataset page: https://huggingface.co/datasets/AdamAxelrod/microscope_pipette_2026-09-02.tabularrobotics10K<n<100K0 likes184 downloads23d agoHugging Face26ShreyasMalwal /MicroG-4M MicroG-4M Dataset This repository stores the entire content of the MicroG-4M dataset itself. For more information and details, including training, evaluation, statistics, and related code, please: Refer to our paper Visit our GitHub And check our fine-tuned models Specification of MicroG-4M "annotation_files" Folder The folder contains all annotation files of the dataset, all stored in CSV format. actions.csv contains all the… See the full description on the dataset page: https://huggingface.co/datasets/ShreyasMalwal/MicroG-4M.documentvideo-classification100K<n<1M0 likes182 downloads22d agoHugging Face27AdamAxelrod /microscope_pipette_2026-09-10This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "action": { "dtype": "float32", "names": [ "joint_1.pos", "joint_2.pos", "joint_3.pos", "joint_4.pos", "joint_5.pos", "joint_6.pos", "precision.state" ], "shape": [… See the full description on the dataset page: https://huggingface.co/datasets/AdamAxelrod/microscope_pipette_2026-09-10.tabularrobotics10K<n<100K0 likes182 downloads15d agoHugging Face28microsoft /tsptabularn<1K2 likes177 downloads1y agoHugging Face29clarayyu22 /gpn-msa-microglia-fulltabular1M<n<10M0 likes170 downloads2y agoHugging Face30microsoft /hnm-search-data HnM Search Dataset Created from Recommendations Dataset This synthetic data-set is created using the recommendations dataset: https://huggingface.co/datasets/einrafh/hnm-fashion-recommendations-data (Use of this dataset is subject to the terms and conditions set forth on the original distribution page. This dataset is intended for non-commercial and research use.) https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/data (DATA ACCESS AND USE:… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/hnm-search-data.imagetext-ranking10M<n<100M2 likes162 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.