datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lm-eval-results-openchat-openchat-3.6-8b-20240522-private
Dataset Card for Evaluation run of openchat/openchat-3.6-8b-20240522
Dataset automatically created during the evaluation run of model openchat/openchat-3.6-8b-20240522
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-openchat-openchat-3.6-8b-20240522-private.ultrachat-sharegpt
UltraChat dataset in ShareGPT format
This is the full UltraChat dataset converted to ShareGPT format.
Pretergeek__OpenChat-3.5-0106_8.99B_40Layers-Appended-details
Dataset Card for Evaluation run of Pretergeek/OpenChat-3.5-0106_8.99B_40Layers-Appended
Dataset automatically created during the evaluation run of model Pretergeek/OpenChat-3.5-0106_8.99B_40Layers-Appended
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pretergeek__OpenChat-3.5-0106_8.99B_40Layers-Appended-details.openchat-spin-slimorca-iter2-datasetcogstack-opengpt-sharegpt
CogStack OpenGPT data in ShareGPT format
openchat_sharegpt4_dataset
openchat/openchat_sharegpt4_dataset
This is a processed version of sharegpt_clean.json from openchat/openchat_sharegpt4_dataset.
Unfortunately, there's not much information about that dataset on Huggingface and GitHub.
Changes:
Corrected turn order
Added language field for detected language
Removed duplicates
Redacted URLs, e-mails, phone numbers
Shuffled
Language code
Number of rows
en
54288
zh
8966
ko
2928
es
2193
fr
1885
ja
1402
de
823
pt
755
ru
507… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/openchat_sharegpt4_dataset.openchat-spin-slimorca-iter0-datasetopenchat__openchat-3.5-0106-details
Dataset Card for Evaluation run of openchat/openchat-3.5-0106
Dataset automatically created during the evaluation run of model openchat/openchat-3.5-0106
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/openchat__openchat-3.5-0106-details.openchat__openchat_3.5-details
Dataset Card for Evaluation run of openchat/openchat_3.5
Dataset automatically created during the evaluation run of model openchat/openchat_3.5
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/openchat__openchat_3.5-details.openchat__openchat-3.6-8b-20240522-details
Dataset Card for Evaluation run of openchat/openchat-3.6-8b-20240522
Dataset automatically created during the evaluation run of model openchat/openchat-3.6-8b-20240522
The dataset is composed of 43 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/openchat__openchat-3.6-8b-20240522-details.openchat__openchat_v3.2_super-details
Dataset Card for Evaluation run of openchat/openchat_v3.2_super
Dataset automatically created during the evaluation run of model openchat/openchat_v3.2_super
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/openchat__openchat_v3.2_super-details.Pretergeek__OpenChat-3.5-0106_32K-PoSE-details
Dataset Card for Evaluation run of Pretergeek/OpenChat-3.5-0106_32K-PoSE
Dataset automatically created during the evaluation run of model Pretergeek/OpenChat-3.5-0106_32K-PoSE
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pretergeek__OpenChat-3.5-0106_32K-PoSE-details.openchat__openchat-3.5-1210-details
Dataset Card for Evaluation run of openchat/openchat-3.5-1210
Dataset automatically created during the evaluation run of model openchat/openchat-3.5-1210
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/openchat__openchat-3.5-1210-details.openchat__openchat_v3.2-details
Dataset Card for Evaluation run of openchat/openchat_v3.2
Dataset automatically created during the evaluation run of model openchat/openchat_v3.2
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/openchat__openchat_v3.2-details.beowolx__CodeNinja-1.0-OpenChat-7B-details
Dataset Card for Evaluation run of beowolx/CodeNinja-1.0-OpenChat-7B
Dataset automatically created during the evaluation run of model beowolx/CodeNinja-1.0-OpenChat-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/beowolx__CodeNinja-1.0-OpenChat-7B-details.openchatgpt-safe-r2I'm too lazy to fill in the dataset card template! Think of it like r1, but after NY - timestamp is XX-01-2023. This is not turbo at this point, it was before 26ths. This must be "alpha", I'm 99% sure.
Has same problems, additional one is missing greetings! "NDA" stuff is missing from this as well!
Pretergeek__openchat-3.5-0106_Rebased_Mistral-7B-v0.2-details
Dataset Card for Evaluation run of Pretergeek/openchat-3.5-0106_Rebased_Mistral-7B-v0.2
Dataset automatically created during the evaluation run of model Pretergeek/openchat-3.5-0106_Rebased_Mistral-7B-v0.2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pretergeek__openchat-3.5-0106_Rebased_Mistral-7B-v0.2-details.openchat_sharegpt_v3_vicuna_formatPretergeek__OpenChat-3.5-0106_10.7B_48Layers-Interleaved-details
Dataset Card for Evaluation run of Pretergeek/OpenChat-3.5-0106_10.7B_48Layers-Interleaved
Dataset automatically created during the evaluation run of model Pretergeek/OpenChat-3.5-0106_10.7B_48Layers-Interleaved
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pretergeek__OpenChat-3.5-0106_10.7B_48Layers-Interleaved-details.openchat-spin-slimorca-iter3-datasetPretergeek__OpenChat-3.5-0106_9.86B_44Layers-Appended-details
Dataset Card for Evaluation run of Pretergeek/OpenChat-3.5-0106_9.86B_44Layers-Appended
Dataset automatically created during the evaluation run of model Pretergeek/OpenChat-3.5-0106_9.86B_44Layers-Appended
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pretergeek__OpenChat-3.5-0106_9.86B_44Layers-Appended-details.Pretergeek__OpenChat-3.5-0106_10.7B_48Layers-Appended-details
Dataset Card for Evaluation run of Pretergeek/OpenChat-3.5-0106_10.7B_48Layers-Appended
Dataset automatically created during the evaluation run of model Pretergeek/OpenChat-3.5-0106_10.7B_48Layers-Appended
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pretergeek__OpenChat-3.5-0106_10.7B_48Layers-Appended-details.KimLan_openchat_sftPretergeek__OpenChat-3.5-0106_8.11B_36Layers-Interleaved-details
Dataset Card for Evaluation run of Pretergeek/OpenChat-3.5-0106_8.11B_36Layers-Interleaved
Dataset automatically created during the evaluation run of model Pretergeek/OpenChat-3.5-0106_8.11B_36Layers-Interleaved
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pretergeek__OpenChat-3.5-0106_8.11B_36Layers-Interleaved-details.Pretergeek__OpenChat-3.5-0106_8.11B_36Layers-Appended-details
Dataset Card for Evaluation run of Pretergeek/OpenChat-3.5-0106_8.11B_36Layers-Appended
Dataset automatically created during the evaluation run of model Pretergeek/OpenChat-3.5-0106_8.11B_36Layers-Appended
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pretergeek__OpenChat-3.5-0106_8.11B_36Layers-Appended-details.Pretergeek__OpenChat-3.5-0106_8.99B_40Layers-Interleaved-details
Dataset Card for Evaluation run of Pretergeek/OpenChat-3.5-0106_8.99B_40Layers-Interleaved
Dataset automatically created during the evaluation run of model Pretergeek/OpenChat-3.5-0106_8.99B_40Layers-Interleaved
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pretergeek__OpenChat-3.5-0106_8.99B_40Layers-Interleaved-details.OpenChat-3k
OpenChat-3k
OpenChat-3k is one of the tiniest chat datasets containing synthetic dialogues in a standard JSON format. It contains 3,708 dialogues generated by Llama.
This dataset can be used for basic question-answering models or potential fine-tunes.
Notice
Usage: This dataset can be modified without the need for credit or attribution.
Format: This dataset is in a standard JSON format.
