datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Ko.SlimOrca원본 데이터셋: Open-Orca/SlimOrca
Ko.WizardLM_evol_instruct_V2_196k이 데이터셋은 자체 구축한 번역기로 WizardLM/WizardLM_evol_instruct_V2_196k을 번역한 데이터셋입니다. 아래 README 페이지도 번역기를 통해 번역되었습니다. 참고 부탁드립니다.
News
🔥 🔥 🔥 [08/11/2023] WizardMath 모델을 출시합니다.
🔥 WizardMath-70B-V1.0 모델은 ChatGPT 3.5, Claude Instant 1 및 PaLM 2 540B 를 포함 하 여 GSM8K에서 일부 폐쇄 소스 LLMs 보다 약간 더 우수 합니다.
🔥 우리의 WizardMath-70B-V1.0 모델은 SOTA 오픈 소스 LLM보다 24.8 포인트 높은 GSM8k Benchmarks에서 81.6 pass@1 을 달성합니다.
🔥 우리의 WizardMath-70B-V1.0 모델은 SOTA 오픈 소스 LLM보다 9.2 포인트 높은 MATH 벤치마크에서 22.7 pass@1 을 달성합니다.… See the full description on the dataset page: https://huggingface.co/datasets/nlp-with-deeplearning/Ko.WizardLM_evol_instruct_V2_196k.Ko.HelpSteer원본 데이터셋: nvidia/HelpSteer
ko.SHP
🚢 Korean Stanford Human Preferences Dataset (Ko.SHP)
이 데이터셋은 자체 구축한 번역기를 활용하여 stanfordnlp/SHP 데이터셋을 번역한 것입니다.
아래의 내용은 해당 번역기로 README 파일을 번역한 것입니다. 참고 부탁드립니다.
If you mention this dataset in a paper, please cite the paper: Understanding Dataset Difficulty with V-Usable Information (ICML 2022).
Summary
SHP는 요리에서 법률 조언에 이르기까지 18가지 다른 주제 영역의 질문/지침에 대한 응답에 대한 385K 집단 인간 선호도 데이터 세트이다.
기본 설정은 다른 응답에 대 한 한 응답의 유용성을 반영 하기 위한 것이며 RLHF 보상 모델 및 NLG 평가 모델 (예: SteamSHP)을 훈련 하는 데… See the full description on the dataset page: https://huggingface.co/datasets/nlp-with-deeplearning/ko.SHP.ko.databricks-dolly-15k원본 데이터셋: databricks/databricks-dolly-15k
deeplearning-tasks-v1
deeplearning-tasks-v1
36 exact tasks for the
deeplearning-env
RL environment, on the topics of Deep Learning (Goodfellow, Bengio & Courville): information
theory, backpropagation, optimisation, the linear algebra used in ML, and numerical stability.
field
meaning
task_id
dl-000 … dl-035
category
information / backprop / optimisation / linalg / numerical
prompt
the question, the units, and the exact shape of the answer
api_description
the fixed network and… See the full description on the dataset page: https://huggingface.co/datasets/eltociear/deeplearning-tasks-v1.Corrector101zhTW
ERNIE for Chinese Spelling Correction 繁體中文
MacBertMaskedLM For Chinese Spelling Correction 繁體中文
wikipedia-zh-20230720-filtered.json 繁體中文
Automatic Corpus Generation-zh 繁體中文
那些自然語言處理 (Natural Language Processing, NLP) 踩的坑 -- 文本糾錯
ko.openhermes원본 데이터셋: teknium/openhermes
ko.lima원본 데이터셋: GAIR/lima
deeplearning-minimind-RL
