CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01microsoft /orca-math-word-problems-200k Dataset Card This dataset contains ~200K grade school math word problems. All the answers in this dataset is generated using Azure GPT4-Turbo. Please refer to Orca-Math: Unlocking the potential of SLMs in Grade School Math for details about the dataset construction. Dataset Sources Repository: microsoft/orca-math-word-problems-200k Paper: Orca-Math: Unlocking the potential of SLMs in Grade School Math Direct Use This dataset has been designed to… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/orca-math-word-problems-200k.textquestion-answering100K<n<1M498 likes25k downloads3y agoHugging Face02microsoft /Updesh_beta 📢 Updesh: Synthetic Multilingual Instruction Tuning Dataset for 13 Indic Languages NOTE: This is an initial $\beta$-release. We plan to release subsequent versions of Updesh with expanded coverage and enhanced quality control. Future iterations will include larger datasets, improved filtering pipelines. Updesh is a large-scale synthetic dataset designed to advance post-training of LLMs for Indic languages. It integrates translated reasoning data and synthesized open-domain… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/Updesh_beta.textquestion-answering1M<n<10M16 likes8.3k downloads8mo agoHugging Face03microsoft /wiki_qa Dataset Card for "wiki_qa" Dataset Summary Wiki Question Answering corpus from Microsoft. The WikiQA corpus is a publicly available set of question and sentence pairs, collected and annotated for research on open-domain question answering. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances default Size of downloaded dataset files: 7.10 MB Size… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/wiki_qa.textquestion-answering10K<n<100K74 likes7.2k downloads3y agoHugging Face04microsoft /orca-agentinstruct-1M-v1 Dataset Card This dataset is a fully synthetic set of instruction pairs where both the prompts and the responses have been synthetically generated, using the AgentInstruct framework. AgentInstruct is an extensible agentic framework for synthetic data generation. This dataset contains ~1 million instruction pairs generated by the AgentInstruct, using only raw text content publicly avialble on the Web as seeds. The data covers different capabilities, such as text editing, creative… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/orca-agentinstruct-1M-v1.textquestion-answering1M<n<10M467 likes2.3k downloads2y agoHugging Face05microsoft /RHELM RHELM: Beyond Static Dialogues Benchmarking Realistic, Heterogeneous, and Evolving Long-Horizon Memory RHELM is a benchmark for evaluating long-horizon memory capabilities in AI assistants. Unlike benchmarks built around static dialogues, RHELM provides realistic, heterogeneous, and temporally evolving memory sources, together with challenging questions that require multi-hop reasoning, temporal synthesis, and hallucination detection. ⚠️ All characters, events, and personal… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/RHELM.textquestion-answering1K<n<10K16 likes1.5k downloads27d agoHugging Face06microsoft /MMLU-CF MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark [📜 Paper] • [🤗 HF Dataset] • [🐱 GitHub] MMLU-CF is a contamination-free and more challenging multiple-choice question benchmark. This dataset contains 10K questions each for the validation set and test set, covering various disciplines. 1. The Motivation of MMLU-CF The open-source nature of these benchmarks and the broad sources of training data for LLMs have inevitably led to… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/MMLU-CF.textquestion-answering10K<n<100K18 likes1.3k downloads2y agoHugging Face07microsoft /XL-DocBench XL-DocBench Evidence-grounded reasoning across hundreds or thousands of pages. Fully verified by 194 human experts. Hongchen Wei1,†,‡, Yuanzhe Wang2,†,‡, Bei Liu2,*, Yifan Yang2, Qi Dai2, Ruichun Ma2, Kai Qiu2, Yunsheng Li2, Dongdong Chen2, Chong Luo2, Zhenzhong Chen1, Baining Guo2 1Wuhan University &nbsp; 2Microsoft &nbsp; †Equal contribution &nbsp; ‡Work done during an internship at MSRA &nbsp; *Project leader Project Page · Paper · Live Leaderboard… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/XL-DocBench.tabularquestion-answering1K<n<10K7 likes982 downloads24d agoHugging Face08microsoft /PEACE PEACE: Empowering Geologic Map Holistic Understanding with MLLMs [Code] [Paper] [Data] Introduction We construct a geologic map benchmark, GeoMap-Bench, to evaluate the performance of MLLMs on geologic map understanding across different abilities, the overview of it is as shown in below Table. Property Description Source USGS(English) CGS(Chinese) Content Image-question pair… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/PEACE.imagequestion-answering1K<n<10K22 likes361 downloads2y agoHugging Face09microsoft /MeetingBank-QA-Summary Dataset Card for MeetingBank-QA-Summary This dataset is introduced in LLMLingua-2 (Pan et al., 2024) and is designed to assess the performance of compressed meeting transcripts on downstream tasks such as question answering (QA) and summarization. It includes 862 meeting transcripts from the test set of meeting transcripts introduced in MeetingBank (Hu et al, 2023) as the context, togeter with QA pairs and summaries that were generated by GPT-4 for each context transcripts.… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/MeetingBank-QA-Summary.textquestion-answeringn<1K17 likes212 downloads2y agoHugging Face10microsoft /LiveDRBench Dataset Card for LiveDRBench: Deep Research as Claim Discovery Arxiv Paper | Hugging Face Dataset | Evaluation Code We propose a formal characterization of the deep research (DR) problem and introduce a new benchmark, LiveDRBench, to evaluate the performance of DR systems. To enable objective evaluation, we define DR using an intermediate output representation that encodes key claims uncovered during search—separating the reasoning challenge from surface-level report generation.… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/LiveDRBench.textquestion-answeringn<1K8 likes181 downloads1y agoHugging Face11openbenchmarks /OB-Inference-Microtasks Inference Microtasks 29 synthetic microtasks with reference answers across meeting-notes lookup, support-ticket triage, and contract-terms extraction. The public set accompanies the OpenBenchmarks Inference Benchmark, which measures single-user delay on short, deliberately easy structured tasks. Dataset contents Configuration Rows Task contract-terms-extraction 10 Extract commercial terms from a technology contract excerpt. meeting-notes-lookup 13… See the full description on the dataset page: https://huggingface.co/datasets/openbenchmarks/OB-Inference-Microtasks.texttext-generationn<1K1 likes129 downloads20d agoHugging Face12ReactiveAI /RealStories-Micro-MRL Dataset Card for ReactiveAI/RealStories-Micro-MRL First synthetic Memory Reinforcement Learning dataset for Proof-of-Concept Reactive Transformer models. Dataset is divided into subsets, used in different Curriculum Stage of MRL training - each subset have different number of follow-up interactions, could use different strategy, and have train and validation splits. Subsets steps-1: ~2300 train (~4600 interactions) / ~340 validation (~680 interactions) - Single-Step… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/RealStories-Micro-MRL.textreinforcement-learning1K<n<10K0 likes105 downloads1y agoHugging Face13MicPie /unpredictable_msdn-microsoft-comThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice10K<n<100K1 likes94 downloads4y agoHugging Face14AmanPriyanshu /reasoning-sft-minimax-microsoft-orca-agentinstruct-1M-v1 MiniMax-M2.5 Reasoning SFT (Orca AgentInstruct 1M v1) Reasoning SFT dataset generated by MiniMaxAI/MiniMax-M2.5 on prompts from the Stratified K-Means Diverse Instruction-Following 100K-1M dataset (Orca AgentInstruct subset). Format Each row has three columns: input — list of dicts [{"role": "...", "content": "..."}, ...] (conversation turns) response — model-generated response with <think> reasoning block source — task category (creative_content, text_modification, rc… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-minimax-microsoft-orca-agentinstruct-1M-v1.texttext-generation100K<n<1M1 likes87 downloads6mo agoHugging Face15LLMTeamAkiyama /cleand_microsoft_rStar-Coder元データ: https://huggingface.co/datasets/microsoft/rStar-Coder データ件数: 269,863 平均トークン数: 11674 最大トークン数: 31,184 合計トークン数: 3,150,447,484 ファイル形式: JSONL ファイルサイズ: 不明 加工内容 synthetic_sftを使用 トークン処理が重たいので、文字数でフィルター seed_question < 6000 generation < 80000 thinkタグ除去 が中途半端なものを除外 トークナイズ処理(速度向上アップデート 繰り返し除去 tabularquestion-answering100K<n<1M0 likes62 downloads1y agoHugging Face16SaveDollars /offline-micro-saas-catalog 📦 SaveDollars.store — Offline Micro SaaS & Autonomous AI Software Catalog This dataset contains structured product metadata, architecture specifications, pricing, and documentation for 96 standalone offline Micro SaaS applications, autonomous AI agent command centers, and business operating systems published by SaveDollars.store. 📊 Dataset Structure (catalog.json) Each record represents a production-ready, subscription-free software package: { "id": 75809… See the full description on the dataset page: https://huggingface.co/datasets/SaveDollars/offline-micro-saas-catalog.tabulartext-generationn<1K1 likes56 downloads12d agoHugging Face17pthinc /BCE-Prettybird-Micro-Standard-v0.0.1 🚀 The Future Standard / Geleceğin Standartı [English] Beyond Raw Data: The Behavioral Revolution The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.1.texttext-generation10K<n<100K1 likes50 downloads7mo agoHugging Face18bio-protocol /bio-faiss-microbiome-v1 bio-faiss-microbiome-v1 A FAISS index + metadata for scientific retrieval Contents index.faiss: FAISS index (cosine w/ inner product). meta.jsonl: one JSON per chunk; fields include chunk_id, paper_id, title, section, subsection, paragraph_index, keywords, boost. Build provenance Chunking: hierarchical (section→paragraph→~380-token chunks, ~15% overlap) Embedder: bio-protocol/scientific-retriever (mean-pooled, L2-normalized) Similarity: cosine via inner… See the full description on the dataset page: https://huggingface.co/datasets/bio-protocol/bio-faiss-microbiome-v1.tabulartext-retrieval10K<n<100K0 likes48 downloads1y agoHugging Face19obekt /obekt-question-answer-reasoning-micro-v0.1 Obekt Micro Reasoning Dataset (v0.1) Dataset Description This is a "micro" dataset containing questions, answers, and reasoning traces. It is generated using the Xiaomi MiMo V2 Flash LLM and is intended for experimental purposes, quick prototyping, and fine-tuning trials where reasoning capability is a focus. Source Model: xiaomi/mimo-v2-flash Contains obekt-question-answer-reasoning-micro-v0.1.csv: The main data file. Columns: question: The input query.… See the full description on the dataset page: https://huggingface.co/datasets/obekt/obekt-question-answer-reasoning-micro-v0.1.texttext-generation10K<n<100K0 likes47 downloads8mo agoHugging Face20pthinc /BCE-Prettybird-Micro-Standard-v0.0.2 🚀 The Future Standard / Geleceğin Standartı [English] Beyond Raw Data: The Behavioral Revolution The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.2.texttext-generation10K<n<100K1 likes47 downloads7mo agoHugging Face215CD-AI /Vietnamese-microsoft-orca-math-word-problems-200k-gg-translatedtexttext-generation100K<n<1M4 likes44 downloads3y agoHugging Face22pthinc /BCE-Prettybird-Micro-Standard-v0.0.4 🚀 The Future Standard / Geleceğin Standartı [English] Beyond Raw Data: The Behavioral Revolution The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.4.texttext-generation10K<n<100K0 likes37 downloads6mo agoHugging Face23pthinc /BCE-Prettybird-Micro-Standard-v0.0.3 🚀 The Future Standard / Geleceğin Standartı [English] Beyond Raw Data: The Behavioral Revolution The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.3.texttext-generation10K<n<100K0 likes30 downloads6mo agoHugging Face24xmarva /psychoanalysis-microtextquestion-answeringn<1K1 likes27 downloads1y agoHugging Face25pthinc /BCE-Prettybird-Micro-Math-v0.1 BCE-Prettybird-Micro-Math-v0.1 10,500 Math Q&A Dataset for Instruction-Based Learning We are excited to introduce a comprehensive math dataset containing 10,500 instruction-based question-answer pairs, designed to support research in mathematical reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to advanced calculus, probability… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Math-v0.1.texttext-classification10K<n<100K0 likes26 downloads6mo agoHugging Face26pthinc /BCE-Prettybird-Micro-Standard-v0.0.5 🚀 The Future Standard / Geleceğin Standartı [English] Beyond Raw Data: The Behavioral Revolution The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.5.texttext-generation10K<n<100K0 likes26 downloads6mo agoHugging Face27pthinc /BCE-Prettybird-Micro-Standard-v0.0.6 🚀 The Future Standard / Geleceğin Standartı [English] Beyond Raw Data: The Behavioral Revolution The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.6.texttext-generation10K<n<100K0 likes26 downloads4mo agoHugging Face28pthinc /turkish_instruct_dataset_micro Turkish Instruct Dataset (Micro) Dataset Description This dataset contains 10169 synthetic samples generated using AI models. It is designed for fine-tuning Turkish LLMs on instruction-following, Q&A, Refusal, Chat, and Logic tasks. Statistics Total Samples: 10169 Categories: Instruction, Q&A, Refusal, Chat, Logic Files turkish_dataset_final_20k.csv: Raw CSV data turkish_dataset_final_20k.json: JSON format turkish_dataset_final_20k.parquet:… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/turkish_instruct_dataset_micro.texttext-generation10K<n<100K0 likes25 downloads8mo agoHugging Face29pthinc /BCE-Prettybird-Micro-Standard-v0.1 Prometech A.Ş. BCE-Prettybird-Micro-Standard-v0.1 🚀 The Future Standard / Geleceğin Standartı [English] Beyond Raw Data: The Behavioral Revolution The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.1.texttext-generation10K<n<100K0 likes19 downloads4mo agoHugging Face30teachaifinance /trade-lifecycle-microstructure-v1 Trade Lifecycle & Market Microstructure Dataset v1 Dataset Summary Trade Lifecycle & Market Microstructure Dataset v1 is a curated, expert-designed dataset focused on market microstructure, trade lifecycle, clearing & settlement, corporate actions, surveillance, and crypto AMM mechanics. The dataset contains 100 high-quality training samples created by a former U.S. equities exchange Market Operations analyst with real-world experience across: U.S. equities… See the full description on the dataset page: https://huggingface.co/datasets/teachaifinance/trade-lifecycle-microstructure-v1.texttext-classificationn<1K0 likes17 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.