datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
m_hellaswag
Multilingual HellaSwag
Dataset Summary
This dataset is a machine translated version of the HellaSwag dataset.
The Icelandic (is) part was translated with Miðeind's Greynir model and Norwegian (nb) was translated with DeepL. The rest of the languages was translated using GPT-3.5-turbo by the University of Oregon, and this part of the dataset was originally uploaded to this Github repository.
hellaswag-trThis Dataset is part of a series of datasets aimed at advancing Turkish LLM Developments by establishing rigid Turkish benchmarks to evaluate the performance of LLM's Produced in the Turkish Language.
Dataset Card for Hellaswag-Turkish
malhajar/hellaswag-turkish is a translated version of hellaswag aimed specifically to be used in the OpenLLMTurkishLeaderboard
This Dataset contains rigid tests extracted from the paper Can a Machine Really Finish Your Sentence? published at ACL2019.… See the full description on the dataset page: https://huggingface.co/datasets/malhajar/hellaswag-tr.m_hellaswag
Multilingual HellaSwag
Dataset Summary
This dataset is a machine translated version of the HellaSwag dataset.
The languages was translated using GPT-3.5-turbo by the University of Oregon, and this part of the dataset was originally uploaded to this Github repository.
The NUS Deep Learning Lab contributed to this effort by standardizing the dataset, ensuring consistent question formatting and alignment across all languages. This standardization enhances cross-linguistic… See the full description on the dataset page: https://huggingface.co/datasets/richmondsin/m_hellaswag.medical-o1-reasoning-SFT
News
[2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data.
[2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1.
[2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an LLM… See the full description on the dataset page: https://huggingface.co/datasets/Hellrabbit/medical-o1-reasoning-SFT.databricks-dolly-15k
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several
of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification,
closed QA, generation, information extraction, open QA, and summarization.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the
Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/Helllloooo7919/databricks-dolly-15k.
