microsoft
rStar-Coder
rStar-Coder Dataset
Project GitHub | Paper
Dataset Description
rStar-Coder is a large-scale competitive code problem dataset containing 418K programming problems, 580K long-reasoning solutions, and rich test cases of varying difficulty levels. This dataset aims to enhance code reasoning capabilities in large language models, particularly in handling competitive code problems.
Experiments on Qwen models (1.5B-14B) across various code reasoning benchmarks demonstrate… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/rStar-Coder.orca-math-word-problems-200k
Dataset Card
This dataset contains ~200K grade school math word problems. All the answers in this dataset is generated using Azure GPT4-Turbo. Please refer to Orca-Math: Unlocking the potential of
SLMs in Grade School Math for details about the dataset construction.
Dataset Sources
Repository: microsoft/orca-math-word-problems-200k
Paper: Orca-Math: Unlocking the potential of
SLMs in Grade School Math
Direct Use
This dataset has been designed to… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/orca-math-word-problems-200k.ms_marco
Dataset Card for "ms_marco"
Dataset Summary
Starting with a paper released at NIPS 2016, MS MARCO is a collection of datasets focused on deep learning in search.
The first dataset was a question answering dataset featuring 100,000 real Bing questions and a human generated answer.
Since then we released a 1,000,000 question dataset, a natural langauge generation dataset, a passage ranking dataset,
keyphrase extraction dataset, crawling dataset, and a conversational search.… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/ms_marco.NOTSOFAR
Introduction
Welcome to the "NOTSOFAR-1: Distant Meeting Transcription with a Single Device" Challenge.
This repo contains the baseline system code for the NOTSOFAR-1 Challenge.
For more information about NOTSOFAR, visit CHiME's official challenge website
Register to participate.
Baseline system description.
Contact us: join the chime-8-notsofar channel on the CHiME Slack, or open a GitHub issue.
📊 Baseline Results on NOTSOFAR dev-set-1
Values are presented in… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/NOTSOFAR.timewarp
Timewarp datasets
This dataset contains molecular dynamics simulation data that was used to train the neural networks in the NeurIPS 2023 paper Timewarp: Transferable Acceleration of Molecular Dynamics by Learning Time-Coarsened Dynamics by Leon Klein, Andrew Y. K. Foong, Tor Erlend Fjelde, Bruno Mlodozeniec, Marc Brockschmidt, Sebastian Nowozin, Frank Noé, and Ryota Tomioka.
Please see the accompanying GitHub repository.
This dataset consists of many molecular dynamics trajectories… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/timewarp.Updesh_beta
📢 Updesh: Synthetic Multilingual Instruction Tuning Dataset for 13 Indic Languages
NOTE: This is an initial $\beta$-release. We plan to release subsequent versions of Updesh with expanded coverage and enhanced quality control. Future iterations will include larger datasets, improved filtering pipelines.
Updesh is a large-scale synthetic dataset designed to advance post-training of LLMs for Indic languages. It integrates translated reasoning data and synthesized open-domain… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/Updesh_beta.
