datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenOrca🐋 The OpenOrca Dataset! 🐋
We are thrilled to announce the release of the OpenOrca dataset!
This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper.
It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers!
Official Models
Mistral-7B-OpenOrca
Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/OpenOrca.SlimOrca
Overview
This is a new curated subset of our OpenOrca data. This release provides an efficient means of reaching performance on-par with using larger slices of our data, while only including ~500k GPT-4 completions.
The key change in this dataset is that we've done an additional pass, using GPT-4 to remove answers which appear wrong based on the human annotations from the FLAN dataset.
This reduces the dataset size to only ~500k entries, allowing training to a similar quality level… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/SlimOrca.SlimOrca-Dedup
Overview
"SlimOrca Dedup" is a deduplicated, unfiltered subset of the SlimOrca dataset, excluding RLHF instances, resulting in 363k unique examples.
Key Features
Removal of RLHF instances.
Deduplication using minhash and Jaccard similarity techniques.
Demo Models
Note: These models were trained on the full SlimOrca dataset, not the deduplicated, unfiltered version.
* https://huggingface.co/openaccess-ai-collective/jackalope-7b
*… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/SlimOrca-Dedup.OpenOrca-tr
Dataset Card for "OpenOrca-tr"
This Dataset is part of a series of datasets aimed at advancing Turkish LLM Developments by establishing rigid Turkish dataset collection to enhance the performance of LLM's Produced in the Turkish Language.
malhajar/orca-tr is a translated version of the OpenOrca and is the first ever SFT dataset in the Turkish Language with more than 2M entries!
Translated by: Mohamad Alhajar
Dataset Summary
The OpenOrca dataset is a collection of… See the full description on the dataset page: https://huggingface.co/datasets/malhajar/OpenOrca-tr.OpenOrca-Traditional-Chinese🐋 OpenOrca-Chinese 数据集!🐋
感謝 Open-Orca/OpenOrca 資料集的發布,為廣大NLP研究人員和開發者帶來了寶貴的資源!
這是一個對 Open-Orca/OpenOrca 資料集中文翻譯的版本,翻譯引擎為 Google 翻譯,希望能為中文 LLM 研究做出一點點貢獻。
Dataset Summary
The OpenOrca dataset is a collection of augmented FLAN Collection data.
Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions.
It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing… See the full description on the dataset page: https://huggingface.co/datasets/lchakkei/OpenOrca-Traditional-Chinese.OpenOrca-Chinese🐋 OpenOrca-Chinese 数据集!🐋
感谢 Open-Orca/OpenOrca 数据集的发布,给广大NLP研究人员和开发者带来了宝贵的资源!
这是一个对 Open-Orca/OpenOrca 数据集中文翻译的版本,翻译引擎为 Google 翻译,希望能给中文 LLM 研究做出一点点贡献。
Dataset Summary
The OpenOrca dataset is a collection of augmented FLAN Collection data.
Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions.
It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing… See the full description on the dataset page: https://huggingface.co/datasets/yys/OpenOrca-Chinese.OpenOrca-ru
OpenOrca-ru
This is translated version of Open-Orca/OpenOrca into Russian.
openorca-chinese-zhtw
Dataset Card for "openorca-chinese-zhtw"
Dataset Summary
The OpenOrca dataset is a collection of augmented FLAN Collection data.
Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions.
It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing generation to expand its scope.
The data is primarily used for training and evaluation in the field of natural… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/openorca-chinese-zhtw.OpenOrca-KO
OpenOrca-KO
OpenOrca dataset 중 약 2만개를 sampling하여 번역한 데이터셋
데이터셋 이용하셔서 모델이나 데이터셋을 만드실 때, 간단한 출처 표기를 해주신다면 연구에 큰 도움이 됩니다😭😭
Dataset inf0
NIV // 1571개
FLAN // 9434개
T0 // 6351개
CoT // 2117개
KoCoT // 2159개
Translation
Using DeepL Pro API. Thanks.
Below is original dataset card
🐋 The OpenOrca Dataset! 🐋
We are thrilled to announce the release of the OpenOrca dataset!
This rich collection of augmented FLAN data aligns, as best as possible, with… See the full description on the dataset page: https://huggingface.co/datasets/kyujinpy/OpenOrca-KO.KOR-OpenOrca-Platypus
KOR-OpenOrca-Platypus
OpenOrca-Ko + KOpen-platypus
데이터셋 이용하셔서 모델이나 데이터셋을 만드실 때, 간단한 출처 표기를 해주신다면 연구에 큰 도움이 됩니다😭😭
KOpen-platpyus
Repo: KOpen-platypus
고품질 한국어 데이터셋
코드와 주석은 그대로 유지하고, 설명 부분만 한국어로 수정
1번과 더불어서, Python, Java, Cpp, xml 등등 결과들은 전부 기존의 데이터 형태로 최대한 보존
단일 숫자와 영어는 본래의 결과 그대로 가져옴
DeepL Pro 번역 결과 중 미완성 변역 결과 직접 수정(예를 들면, '[...]'가 포함되어 있음)
DeepL Pro 번역 결과가 본래의 데이터에 비해 글자수가 50% 이하로 낮으면, 번역 결과 수정
번역하고자 하는 글자수가 1500자 이상일 경우, API로 변경해서 번역
고유명사는 최대한 유지함
Post-processing… See the full description on the dataset page: https://huggingface.co/datasets/kyujinpy/KOR-OpenOrca-Platypus.KOR-OpenOrca-Platypus-v3
KOR-OpenOrca-Platypus-v3
KOR-OpenOrca-Platypus 데이터셋에서 수작업으로 번역 오류 200건 이상을 고친 데이터셋.
데이터셋 이용하셔서 모델이나 데이터셋을 만드실 때, 간단한 출처 표기를 해주신다면 연구에 큰 도움이 됩니다😭😭
KOpen-platpyus
Repo: KOpen-platypus
고품질 한국어 데이터셋
코드와 주석은 그대로 유지하고, 설명 부분만 한국어로 수정
1번과 더불어서, Python, Java, Cpp, xml 등등 결과들은 전부 기존의 데이터 형태로 최대한 보존
단일 숫자와 영어는 본래의 결과 그대로 가져옴
DeepL Pro 번역 결과 중 미완성 변역 결과 직접 수정(예를 들면, '[...]'가 포함되어 있음)
DeepL Pro 번역 결과가 본래의 데이터에 비해 글자수가 50% 이하로 낮으면, 번역 결과 수정
번역하고자 하는 글자수가 1500자 이상일 경우, API로… See the full description on the dataset page: https://huggingface.co/datasets/kyujinpy/KOR-OpenOrca-Platypus-v3.OpenOrca-zh-20k
Datsetcard for 'OpenOrca-zh-20k'
This is the Chinese version of Open-Orca/OpenOrca from Azure99/blossom-orca-v3.
Compared to Azure99/blossom-orca-v3:
This dataset extracts all Chinese blossom-orca-v3 samples (around 20K) into a separate zh split.
All samples are formatted in the ocra format with an optional system role in the first round.
Instead of using a 1:1 En-Zh ratio as in blossom-orca-v3, this dataset contains 200K GPT-4 generated English samples from OpenOrca in the en… See the full description on the dataset page: https://huggingface.co/datasets/wenbopan/OpenOrca-zh-20k.OpenOrca🐋 The OpenOrca Dataset! 🐋
We are thrilled to announce the release of the OpenOrca dataset!
This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper.
It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers!
Official Models
Mistral-7B-OpenOrca
Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/polinaeterna/OpenOrca.ChatML-OpenOrcaOpen-Orca/OpenOrca in ChatML format, ready to use in HuggingFace TRL's SFT Trainer.
Python code used for conversion:
from datasets import load_dataset
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("Felladrin/Minueza-32M-Base")
dataset = load_dataset("Open-Orca/OpenOrca", split="train")
def format(columns):
messages = []
system_prompt = columns["system_prompt"].strip()
if system_prompt:
messages.append({
"role": "system"… See the full description on the dataset page: https://huggingface.co/datasets/Felladrin/ChatML-OpenOrca.OpenOrca-Ko-En
OpenOrca-Ko-En
kyujinpy/OpenOrca-KO와 Open-Orca/OpenOrca를 공통된 데이터만 필터링하고 합친 데이터셋입니다.
컬럼은 기존 OpenOrca에 맞춰서 system_prompt_{ko/en}, question_{ko/en}, response_{ko/en} 으로 변경하였습니다.
중복된 id를 제거하여 데이터수가 일부 감소하였습니다.
데이터셋을 만드는데 사용한 스크립트입니다.
데이터셋 이용하셔서 모델이나 데이터셋을 만드실 때, 이 데이터셋뿐만 아니라 위 데이터셋도 함께 출처표기를 해주셨으면 합니다.
Dataset inf0
NIV // 1551개(OpenOrca-KO: 1571개)
FLAN // 9338개(OpenOrca-KO: 9434개)
T0 // 6303개(OpenOrca-KO: 6351개)
CoT // 2092개(OpenOrca-KO: 2117개)
KoCoT // 2159개… See the full description on the dataset page: https://huggingface.co/datasets/appleparan/OpenOrca-Ko-En.OpenOrca-gugugo-ko
OpenOrca 한국어 번역 데이터셋
Gugugo-koen-7B-V1.1을 이용하여 OpenOrca데이터셋을 번역하고 있습니다.
번역 진행상황은 아래를 참고해 주십시오.
진행상황
GPT4 생성물 약 100만 개 중 약 64만 개 번역완료
GPT3.5 생성물 약 350만 개 중 약 159만 개 번역완료
데이터셋 사용 후 출처표기는 제작자에게 큰 힘이 됩니다.
Original dataset card: OpenOrca
🐋 The OpenOrca Dataset! 🐋
We are thrilled to announce the release of the OpenOrca dataset!
This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper.
It has… See the full description on the dataset page: https://huggingface.co/datasets/squarelike/OpenOrca-gugugo-ko.Vietnamese-Openorca-Multiplechoice-gg-translatedKOR-OpenOrca-Platypus-v2
KOR-OpenOrca-Platypus-v2
KOR-OpenOrca-Platypus 데이터셋에서 수작업으로 번역 오류 200건 이상을 고친 데이터셋.
데이터셋 이용하셔서 모델이나 데이터셋을 만드실 때, 간단한 출처 표기를 해주신다면 연구에 큰 도움이 됩니다😭😭
KOpen-platpyus
Repo: KOpen-platypus
고품질 한국어 데이터셋
코드와 주석은 그대로 유지하고, 설명 부분만 한국어로 수정
1번과 더불어서, Python, Java, Cpp, xml 등등 결과들은 전부 기존의 데이터 형태로 최대한 보존
단일 숫자와 영어는 본래의 결과 그대로 가져옴
DeepL Pro 번역 결과 중 미완성 변역 결과 직접 수정(예를 들면, '[...]'가 포함되어 있음)
DeepL Pro 번역 결과가 본래의 데이터에 비해 글자수가 50% 이하로 낮으면, 번역 결과 수정
번역하고자 하는 글자수가 1500자 이상일 경우, API로… See the full description on the dataset page: https://huggingface.co/datasets/kyujinpy/KOR-OpenOrca-Platypus-v2.OpenOrca🐋 The OpenOrca Dataset! 🐋
We are thrilled to announce the release of the OpenOrca dataset!
This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper.
It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers!
Official Models
Mistral-7B-OpenOrca
Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/mainbrains/OpenOrca.OpenOrca-Top5percent🐋 The OpenOrca-Top5Percent Dataset! 🐋
We are excited to introduce the OpenOrca-Top5Percent dataset, a refined version of the original OpenOrca dataset. This dataset contains only those entries which utilize the top 5% most frequently used words in the OpenOrca dataset, aiming to focus on high-frequency vocabulary for various NLP tasks.
Dataset Summary
The OpenOrca-Top5Percent dataset is a curated subset of the augmented FLAN Collection data, focusing specifically on entries that… See the full description on the dataset page: https://huggingface.co/datasets/dynopii/OpenOrca-Top5percent.OpenOrca🐋 The OpenOrca Dataset! 🐋
We are thrilled to announce the release of the OpenOrca dataset!
This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper.
It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers!
Official Models
Mistral-7B-OpenOrca
Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/Preeetam/OpenOrca.OpenOrca-Open-Orca🐋 The OpenOrca Dataset! 🐋
We are thrilled to announce the release of the OpenOrca dataset!
This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper.
It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers!
Official Models
Mistral-7B-OpenOrca
Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/GenRM/OpenOrca-Open-Orca.OpenOrca🐋 The OpenOrca Dataset! 🐋
We are thrilled to announce the release of the OpenOrca dataset!
This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper.
It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers!
Official Models
Mistral-7B-OpenOrca
Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/KIZZ2006/OpenOrca.OpenOrcaNo-15k🐋 The OpenOrca Dataset Norwegian! 🐋
This is a subset of 15000 rows of the OpenOrca dataset, translated into Norwegian.
Translation is done with Amazon Translate, and is provided by Ruter as an artifact from Ruter AI Lab.
Dataset structure
The dataset is structured in the following way:
{
"instruction": "Norwegian instruction",
"input": "Norwegian input",
"output": "Norwegian output",
"instruction_en": "English instruction",
"input_en": "English input"… See the full description on the dataset page: https://huggingface.co/datasets/RuterNorway/OpenOrcaNo-15k.openorca-zht
Dataset Card for "openorca-chinese-zhtw"
Dataset Summary
The OpenOrca dataset is a collection of augmented FLAN Collection data.
Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions.
It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing generation to expand its scope.
The data is primarily used for training and evaluation in the field of… See the full description on the dataset page: https://huggingface.co/datasets/shutajb/openorca-zht.1M-OpenOrca_beEn/Be
🐋 The Belarusian OpenOrca Dataset! 🐋
Belarusian OpenOrca dataset - is rich collection of augmented FLAN data aligns, that translated in belarusian language.
That dataset should help training LLM in belarusian language and should help on other NLP tasks.
This dataset have 2 version:
~1M GPT-4 completions (Now translating)
~3.2M GPT-3.5 completions (Can be translated in future)
Data Fields
The fields are:
'id', a unique numbered identifier which includes one of 'niv'… See the full description on the dataset page: https://huggingface.co/datasets/WiNE-iNEFF/1M-OpenOrca_be.OpenOrca_35k
Dataset Card for "OpenOrca_35k"
The first 35k examples from Open-Orca/OpenOrca
OpenOrca🐋 The OpenOrca Dataset! 🐋
We are thrilled to announce the release of the OpenOrca dataset!
This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper.
It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers!
Official Models
Mistral-7B-OpenOrca
Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/villagerad/OpenOrca.open-orca-slimorca-deduped-cleaned-corrected-for-pascal-txtThis is a modified version of the slimorca-deduped-cleaned-corrected dataset.
It contains English only characters.
Open Orca Slim for Pascal Developers is a subset of the original Open Orca dataset .
Open Orca Slim for Pascal Developers dataset was created with:
from datasets import load_dataset
# Coded by Gemini
def biggest_char_code(input_string):
"""
Returns the largest character code in a string.
"""
if not input_string:
return None # Handle empty string case
largest_code… See the full description on the dataset page: https://huggingface.co/datasets/schuler/open-orca-slimorca-deduped-cleaned-corrected-for-pascal-txt.
