CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01microsoft /SCBench SCBench [Paper] [Code] [Project Page] SCBench (SharedContextBench) is a comprehensive benchmark to evaluate efficient long-context methods in a KV cache-centric perspective, analyzing their performance across the full KV cache lifecycle (generation, compression, retrieval, and loading) in real-world scenarios where context memory (KV cache) is shared and reused across multiple requests. 🎯 Quick Start Load Data You can download and load the SCBench data… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/SCBench.tabularn<1K10 likes2.2k downloads2y agoHugging Face02microsoft /Taskbench TaskBench: Benchmarking Large Language Models for Task Automation Introduction TaskBench is a benchmark for evaluating large language models (LLMs) on task automation. Task automation can be formulated into three critical stages: task decomposition, tool invocation, and parameter prediction. This complexity makes data collection and evaluation more challenging compared to common NLP tasks. To address this challenge, we propose a comprehensive evaluation framework… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/Taskbench.tabular10K<n<100K38 likes1.3k downloads2y agoHugging Face03microsoft /XL-DocBench XL-DocBench Evidence-grounded reasoning across hundreds or thousands of pages. Fully verified by 194 human experts. Hongchen Wei1,†,‡, Yuanzhe Wang2,†,‡, Bei Liu2,*, Yifan Yang2, Qi Dai2, Ruichun Ma2, Kai Qiu2, Yunsheng Li2, Dongdong Chen2, Chong Luo2, Zhenzhong Chen1, Baining Guo2 1Wuhan University &nbsp; 2Microsoft &nbsp; †Equal contribution &nbsp; ‡Work done during an internship at MSRA &nbsp; *Project leader Project Page · Paper · Live Leaderboard… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/XL-DocBench.tabularquestion-answering1K<n<10K7 likes932 downloads20d agoHugging Face04microsoft /kitab Overview 🕮 KITAB is a challenging dataset and a dynamic data collection approach for testing abilities of Large Language Models (LLMs) in answering information retrieval queries with constraint filters. A filtering query with constraints can be of the form "List all books written by Toni Morrison that were published between 1970-1980". The dataset was originally contributed by the paper "KITAB: Evaluating LLMs on Constraint Satisfaction for Information Retrieval" Marah I Abdin… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/kitab.tabular10K<n<100K13 likes853 downloads3y agoHugging Face05microsoft /bing_coronavirus_query_set Dataset Card for BingCoronavirusQuerySet Dataset Summary Please note that you can specify the start and end date of the data. You can get start and end dates from here: https://github.com/microsoft/BingCoronavirusQuerySet/tree/master/data/2020 example: load_dataset("bing_coronavirus_query_set", queries_by="state", start_date="2020-09-01", end_date="2020-09-30") You can also load the data by country by using queries_by="country". Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/bing_coronavirus_query_set.tabulartext-classification100K<n<1M1 likes469 downloads3y agoHugging Face06microsoft /MuseVLA-dataset MuseVLA Dataset Multi-modal robot manipulation dataset with synchronized RGB, depth, acoustic, thermal, and radar streams. Released as two parts (dataset_01/, dataset_02/) sharing the same per-episode layout. Together they cover ~1400 episodes across 11 instructions (towel / clothes / box / item / drink manipulation). Per-episode contents {episode_name}/ ├── video.mp4 # RGB, 1280×720, 30 fps ├── mask/video.mp4 #… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/MuseVLA-dataset.tabularrobotics1K<n<10K2 likes453 downloads1mo agoHugging Face07microsoft /msr-acc-tae25 Microsoft Research - Accurate Chemistry Collection: Total Atomization Energies Description The Microsoft Research Accurate Chemistry Collection (MSR-ACC) provides a collection of accurate coupled cluster labels for training machine learning functionals. MSR-ACC/TAE25 comprising 73,040 total atomization energies at the CCSD(T)/CBS level obtained with the W1-F12 thermochemical protocol. The dataset is constructed to exhaustively cover the chemical space of closed-shell… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/msr-acc-tae25.tabular10K<n<100K8 likes357 downloads5mo agoHugging Face08microsoft /mediflow MediFlow A large-scale synthetic instruction dataset of 2.5M rows (~700k unique instructions) for clinical natural language processing covering 14 task types and 98 fine-grained input clinical documents. t-SNE 2D Plot of MediFlow Embeddings by Task Types Dataset Splits mediflow: 2.5M instruction data for SFT alignment. mediflow_dpo: ~135k top-quality instructions with GPT-4o generated rejected_output for DPO alignment. Main Columns instruction:… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/mediflow.tabulartext-generation1M<n<10M53 likes298 downloads9mo agoHugging Face09microsoft /WildFeedback Dataset Card for WildFeedback WildFeedback is a preference dataset constructed from real-world user interactions with ChatGPT. Unlike synthetic datasets that rely solely on AI-generated rankings, WildFeedback captures authentic human preferences through naturally occurring user feedback signals in conversation. The dataset is designed to improve the alignment of large language models (LLMs) with actual human values by leveraging direct user input. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/WildFeedback.tabulartext-generation1M<n<10M16 likes265 downloads1y agoHugging Face10microsoft /WorkflowPerturb WorkflowPerturb — Dataset Artifact Companion data for the EMNLP 2026 Industry Track paper “WorkflowPerturb: Calibrated Stress Tests for Evaluating Multi-Agent Workflow Metrics.” Canonical location: https://huggingface.co/datasets/microsoft/WorkflowPerturbPaper: https://arxiv.org/abs/2602.17990 This release is the complete WorkflowPerturb benchmark plus documentation. It is self-contained: the CSVs carry every golden workflow, every perturbed variant, and all shipped pre-computed… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/WorkflowPerturb.tabulartext-generation10K<n<100K3 likes248 downloads10d agoHugging Face11microsoft /benchpress-score-matrix BenchPress Score Matrix This dataset contains the public model-by-benchmark score matrix used by BenchPress. The release includes the lossless audited JSON, benchmark cost evidence, flat model and benchmark metadata, one row per observed score, and the paper-canonical dense subset used in the BenchPress experiments. The source repository is microsoft/benchpress. Canonical artifacts data/llm_benchmark_data.json is the authoritative rich score-matrix artifact. It… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/benchpress-score-matrix.tabulartabular-regressionn<1K2 likes242 downloads1mo agoHugging Face12microsoft /PatientSafetyBench Disclaimer The synthetic prompts may contain offensive, discriminatory, or harmful language. These fake prompts also mention topics that are not based on the scientific consensus at all.These prompts are included solely for the purpose of evaluating safety behavior of language models. ⚠️ Disclaimer: The presence of such prompts does not reflect the views, values, or positions of the authors, their institutions, or any affiliated organizations. They are provided exclusively for… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/PatientSafetyBench.tabulartext-generationn<1K10 likes164 downloads5mo agoHugging Face13microsoft /hnm-search-data HnM Search Dataset Created from Recommendations Dataset This synthetic data-set is created using the recommendations dataset: https://huggingface.co/datasets/einrafh/hnm-fashion-recommendations-data (Use of this dataset is subject to the terms and conditions set forth on the original distribution page. This dataset is intended for non-commercial and research use.) https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/data (DATA ACCESS AND USE:… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/hnm-search-data.imagetext-ranking10M<n<100M2 likes162 downloads7mo agoHugging Face14microsoft /CoSAlign-Train CoSAlign-Train: A Large-Scale Synthetic Training Dataset for Controllable Safety Alignment Paper: Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements, published at ICLR 2025. Purpose: Training dataset for controllable safety alignment (CoSA) of large language models (LLMs), facilitating fine-grained inference-time adaptation to diverse safety requirements. Description: CoSAlign-Train is a large-scale, synthetic preference dataset designed for… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/CoSAlign-Train.tabular100K<n<1M4 likes129 downloads1y agoHugging Face15microsoft /tsptabularn<1K2 likes114 downloads1y agoHugging Face16LLMTeamAkiyama /cleand_microsoft_rStar-Coder元データ: https://huggingface.co/datasets/microsoft/rStar-Coder データ件数: 269,863 平均トークン数: 11674 最大トークン数: 31,184 合計トークン数: 3,150,447,484 ファイル形式: JSONL ファイルサイズ: 不明 加工内容 synthetic_sftを使用 トークン処理が重たいので、文字数でフィルター seed_question < 6000 generation < 80000 thinkタグ除去 が中途半端なものを除外 トークナイズ処理(速度向上アップデート 繰り返し除去 tabularquestion-answering100K<n<1M0 likes71 downloads1y agoHugging Face17toksuitebackup /microsoft-Phi-3-mini-4k-instruct-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M1 likes65 downloads10mo agoHugging Face18open-llm-leaderboard /microsoft__phi-4-detailsgated Dataset Card for Evaluation run of microsoft/phi-4 Dataset automatically created during the evaluation run of model microsoft/phi-4 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__phi-4-details.tabular10K<n<100K0 likes62 downloads2y agoHugging Face19microsoft /sattabularn<1K2 likes54 downloads1y agoHugging Face20open-llm-leaderboard /microsoft__phi-2-detailsgated Dataset Card for Evaluation run of microsoft/phi-2 Dataset automatically created during the evaluation run of model microsoft/phi-2 The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__phi-2-details.tabular10K<n<100K0 likes49 downloads2y agoHugging Face21open-llm-leaderboard /microsoft__Phi-3-mini-4k-instruct-detailsgated Dataset Card for Evaluation run of microsoft/Phi-3-mini-4k-instruct Dataset automatically created during the evaluation run of model microsoft/Phi-3-mini-4k-instruct The dataset is composed of 73 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3-mini-4k-instruct-details.tabular10K<n<100K0 likes46 downloads2y agoHugging Face22open-llm-leaderboard /microsoft__phi-1_5-detailsgated Dataset Card for Evaluation run of microsoft/phi-1_5 Dataset automatically created during the evaluation run of model microsoft/phi-1_5 The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__phi-1_5-details.tabular10K<n<100K0 likes38 downloads2y agoHugging Face23open-llm-leaderboard /microsoft__Phi-3-medium-4k-instruct-detailsgated Dataset Card for Evaluation run of microsoft/Phi-3-medium-4k-instruct Dataset automatically created during the evaluation run of model microsoft/Phi-3-medium-4k-instruct The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3-medium-4k-instruct-details.tabular10K<n<100K0 likes38 downloads2y agoHugging Face24open-llm-leaderboard /microsoft__Phi-3-mini-128k-instruct-detailsgated Dataset Card for Evaluation run of microsoft/Phi-3-mini-128k-instruct Dataset automatically created during the evaluation run of model microsoft/Phi-3-mini-128k-instruct The dataset is composed of 40 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3-mini-128k-instruct-details.tabular10K<n<100K0 likes37 downloads2y agoHugging Face25emilpartow /reddit_finance_posts_apple-tesla-microsoft Reddit Finance Posts Dataset (Apple, Tesla, Microsoft) This dataset contains 12046 Reddit Posts collected from 20 finance-related subreddits via the Reddit API using the keywords Apple, Tesla, and Microsoft. The included subreddits are: stocks, wallstreetbets, investing, StockMarket, options, RobinHood, pennystocks, SecurityAnalysis, personalfinance, Dividends, CryptoCurrency, CryptoMarkets, ETFs, FinancialIndependence, ValueInvesting, quant, algotrading, forex, economy, Superstonk… See the full description on the dataset page: https://huggingface.co/datasets/emilpartow/reddit_finance_posts_apple-tesla-microsoft.tabularfeature-extraction10K<n<100K4 likes37 downloads1y agoHugging Face26open-llm-leaderboard /microsoft__Phi-4-mini-instruct-detailsgated Dataset Card for Evaluation run of microsoft/Phi-4-mini-instruct Dataset automatically created during the evaluation run of model microsoft/Phi-4-mini-instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-4-mini-instruct-details.tabular10K<n<100K0 likes30 downloads2y agoHugging Face27open-llm-leaderboard /microsoft__Phi-3.5-mini-instruct-detailsgated Dataset Card for Evaluation run of microsoft/Phi-3.5-mini-instruct Dataset automatically created during the evaluation run of model microsoft/Phi-3.5-mini-instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3.5-mini-instruct-details.tabular10K<n<100K0 likes27 downloads2y agoHugging Face28open-llm-leaderboard /microsoft__Phi-3.5-MoE-instruct-detailsgated Dataset Card for Evaluation run of microsoft/Phi-3.5-MoE-instruct Dataset automatically created during the evaluation run of model microsoft/Phi-3.5-MoE-instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3.5-MoE-instruct-details.tabular10K<n<100K0 likes27 downloads2y agoHugging Face29open-llm-leaderboard /microsoft__Phi-3-medium-128k-instruct-detailsgated Dataset Card for Evaluation run of microsoft/Phi-3-medium-128k-instruct Dataset automatically created during the evaluation run of model microsoft/Phi-3-medium-128k-instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3-medium-128k-instruct-details.tabular10K<n<100K0 likes26 downloads2y agoHugging Face30open-llm-leaderboard /microsoft__phi-1-detailsgated Dataset Card for Evaluation run of microsoft/phi-1 Dataset automatically created during the evaluation run of model microsoft/phi-1 The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__phi-1-details.tabular10K<n<100K0 likes25 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.