datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SlimOrca
Overview
This is a new curated subset of our OpenOrca data. This release provides an efficient means of reaching performance on-par with using larger slices of our data, while only including ~500k GPT-4 completions.
The key change in this dataset is that we've done an additional pass, using GPT-4 to remove answers which appear wrong based on the human annotations from the FLAN dataset.
This reduces the dataset size to only ~500k entries, allowing training to a similar quality level… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/SlimOrca.OpenOrca-Viet
🇻🇳 Vietnamese OpenOrca is here 🐋
Dive into the Vietnamese linguistic landscape with OpenOrca, a cutting-edge dataset crafted through a pioneering partnership between Virtual Interactive and Alignment Lab AI. Drawing inspiration and methodology from the renowned Orca paper, we've expanded our horizons to distill knowledge from a more eclectic mix of leading LLMs including GPT-4, PaLM-2, and Claude. Our vision with this dataset is to fuel research and development that will… See the full description on the dataset page: https://huggingface.co/datasets/vilm/OpenOrca-Viet.OpenOrca-Step-by-step-reasoningThis work was performed to help models with reasoning. I developed it working on my Cinder model, a STEM q and a model.
Modified OpenORCA Step-by-Step Reasoning Dataset Overview
The Modified OpenORCA Step-by-Step Reasoning Dataset represents a groundbreaking resource in the field of artificial intelligence, specifically designed to enhance the reasoning capabilities of AI models. This unique dataset is the result of a meticulous process of sorting, selecting, and altering dialogues from the… See the full description on the dataset page: https://huggingface.co/datasets/Josephgflowers/OpenOrca-Step-by-step-reasoning.openorca-multiplechoice-10kA 10k subset of OpenOrca dataset, focusing on multiple choice questions.
Credit to Tian Xia.
OpenOrcaJapaneseOpenOrcaデータセットの日本語翻訳版です
https://huggingface.co/datasets/Open-Orca/OpenOrca
現在翻訳作業が続行中で、OpenOrca全体の1/5程度の翻訳が終わった状態でひとまず公開します。商用利用可能です。
OpenOrca-gugugo-ko
OpenOrca 한국어 번역 데이터셋
Gugugo-koen-7B-V1.1을 이용하여 OpenOrca데이터셋을 번역하고 있습니다.
번역 진행상황은 아래를 참고해 주십시오.
진행상황
GPT4 생성물 약 100만 개 중 약 64만 개 번역완료
GPT3.5 생성물 약 350만 개 중 약 159만 개 번역완료
데이터셋 사용 후 출처표기는 제작자에게 큰 힘이 됩니다.
Original dataset card: OpenOrca
🐋 The OpenOrca Dataset! 🐋
We are thrilled to announce the release of the OpenOrca dataset!
This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper.
It has… See the full description on the dataset page: https://huggingface.co/datasets/squarelike/OpenOrca-gugugo-ko.Vietnamese-Openorca-Multiplechoice-gg-translatedOpenOrca-test-jpTrainデータセット日本語訳は、shumpei2525さんのレポジトリにあります。
https://huggingface.co/datasets/shumpei2525/OpenOrca-train-ja
Here is the dataset:shumpei2525/OpenOrca-train-ja
2023年7月14日時点
Open Orcaが公開したデータセット(GPT-3.5版のtestデータのみ)の翻訳です.
翻訳が失敗しているデータが存在します。
翻訳した結果、意味のないタスクとなっている場合があります。
ライセンスはmit引継ぎとしましたが、
商用利用に関しては、OpenAIの規約がよくわからないので注意してください。
以下は、Open OrcaデータセットのREADME.mdの日本語訳です。
🐋 Open Orca データセット! 🐋
Open Orca… See the full description on the dataset page: https://huggingface.co/datasets/pyutax68/OpenOrca-test-jp.OpenOrca-Traditional-Chinese_structgpt4-1m-orca-embeddingsOpen-Orca__Mistral-7B-OpenOrca-details
Dataset Card for Evaluation run of Open-Orca/Mistral-7B-OpenOrca
Dataset automatically created during the evaluation run of model Open-Orca/Mistral-7B-OpenOrca
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Open-Orca__Mistral-7B-OpenOrca-details.OpenOrca-tr-1-million-sharegpt
OpenOrca-tr-1-million-sharegpt
This dataset is the Turkish version of the Open-Orca/OpenOrca dataset.
This dataset consists of 1 million selected rows, converted to ShareGPT format for compatibility.
openorca-multiplechoice-5k-comparisonsA subset of beaugogh/openorca-multiplechoice-10k, where model responses are added as the "rejected" responses.
The model used here is beaugogh/Llama2-7b-openorca-mc-v2.
OpenOrca-ruОригинал: d0rj/OpenOrca-ru
Здесь записаны 25 тысяч строк, которые отфильтровали и убрали часть "мусора".
openorca_persianOpenOrca-2-fact-viOpenOrca-translate-openQAOpenOrca-movieplot-viOpenOrca-predict-people-action-viOpenOrca-describe-viOpenOrca-solution-for-a-goal-viOpenOrca-conclusion-condition-viOpenOrcaopenorca-alpaca-15kopenorca-alpaca-50kopenorca_data.jsonjosephgflowers_openorca_extractedOpenOrca-300k
