datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenOrca-tr
Includes a part of OpenOrca dataset in Turkish language
The Subset of OpenOrca dataset in turkish language comprises 798350 pairs of questions and answers in Turkish,
predominantly translated from English using Google Translate.
Wherever possible, specific terminology and unique names were retained unchanged in the translation process.
Feel free to submit pull requests to enhance the quality of the dataset.
Contact: https://www.linkedin.com/in/ugur-cekmez/
OpenOrca-Top5percent🐋 The OpenOrca-Top5Percent Dataset! 🐋
We are excited to introduce the OpenOrca-Top5Percent dataset, a refined version of the original OpenOrca dataset. This dataset contains only those entries which utilize the top 5% most frequently used words in the OpenOrca dataset, aiming to focus on high-frequency vocabulary for various NLP tasks.
Dataset Summary
The OpenOrca-Top5Percent dataset is a curated subset of the augmented FLAN Collection data, focusing specifically on entries that… See the full description on the dataset page: https://huggingface.co/datasets/dynopii/OpenOrca-Top5percent.OpenOrca-train-ja
OpenOrca-train-ja
This dataset is a translation of OpenOrca into Japanese. It is based on the output data from GPT-3.5 and GPT-4. Please feel free to use it as you wish.
* There are a few mistakes observed in the translation task. It might be better to exclude the translation task from use.
Since I'm not entirely clear on OpenAI's terms of service, please be cautious when using it for commercial purposes. There may be exceptions for non-commercial use.… See the full description on the dataset page: https://huggingface.co/datasets/shumpei2525/OpenOrca-train-ja.sample-OpenOrcaopenorca7B-cotOpenOrca_cleaned_kor_linkbricks_single_dataset_with_prompt_text_huggingfaceopenorca1open_Orca_preprocessedopen_orcaOPEN_ORCA_MODIFIED_DATAopenorca-10k-subset-llama-2open_orca_tokenizaton_and
