datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-instruct-gptj-pairwiseopenai_summarize_comparisons_relabel_GPTJdetails_digitous__Skegma-GPTJ
Dataset Card for Evaluation run of digitous/Skegma-GPTJ
Dataset Summary
Dataset automatically created during the evaluation run of model digitous/Skegma-GPTJ on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_digitous__Skegma-GPTJ.details_digitous__Janin-GPTJ
Dataset Card for Evaluation run of digitous/Janin-GPTJ
Dataset Summary
Dataset automatically created during the evaluation run of model digitous/Janin-GPTJ on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_digitous__Janin-GPTJ.details_digitous__Adventien-GPTJ
Dataset Card for Evaluation run of digitous/Adventien-GPTJ
Dataset Summary
Dataset automatically created during the evaluation run of model digitous/Adventien-GPTJ on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_digitous__Adventien-GPTJ.details_digitous__Javelin-GPTJ
Dataset Card for Evaluation run of digitous/Javelin-GPTJ
Dataset Summary
Dataset automatically created during the evaluation run of model digitous/Javelin-GPTJ on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_digitous__Javelin-GPTJ.gpt-jokessynthetic-instruct-gptj-pairwise-ru
Dataset Card for "synthetic-instruct-gptj-pairwise-ru"
This is translated version of Dahoas/synthetic-instruct-gptj-pairwise dataset into Russian.
synthetic-instruct-gptj_pairwisesynthetic-instruct-gptj-pairwise
Dataset Card for "synthetic-instruct-gptj-pairwise"
More Information needed
synthetic-instruct-gptj-pairwise-ja
Dahoas/synthetic-instruct-gptj-pairwise-ja
Dahoas/synthetic-instruct-gptj-pairwiseの和訳
sft-gptj-synthetic-prompt-responsescounterfact-filtered-gptj6b
Dataset Card for "counterfact-filtered-gptj6b"
This dataset is a subset of azhx/counterfact-easy, however it was filtered based on a heuristic that was used to determine whether the knowledge in each row is actually known by the GPT-J-6B model
The heuristic is as follows:
For each prompt in the original counterfact dataset used by ROME, we use GPT-J-6B to generate n=5 completions to a max generated token length of 30.
If the completion contains the answer that is… See the full description on the dataset page: https://huggingface.co/datasets/azhx/counterfact-filtered-gptj6b.synthetic_gptj_paraphrasedhh-rrhf-dahoas-gptj-rm
Dataset Card for "hh-rrhf-dahoas-gptj-rm"
More Information needed
data-guided-scp-gptj-litGPT_judgement_base_vs_sft_dpogpt-j-oasst1-es
OpenAssistant Conversations Spanish Dataset (OASST1-es) for GPT-j
Dataset Summary
Subset of the original OpenAssistant Conversations Dataset (OASST).
Filtered by lang=es.
Formatted according to the "instruction - output" pattern.
Select the best ranked output (Some instructions have multiple outputs ranked by humans).
Select only the first level of the tree conversation.
Dataset Structure
The dataset has 3909 rows of tuples (instructions and outputs).
details_digitous__Javalion-GPTJ
Dataset Card for Evaluation run of digitous/Javalion-GPTJ
Dataset Summary
Dataset automatically created during the evaluation run of model digitous/Javalion-GPTJ on the Open LLM Leaderboard.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_digitous__Javalion-GPTJ.GPT_judgement_sftgptjodoo100hh-rrhf-dahoas-gptj-rm-25k
Dataset Card for "hh-rrhf-dahoas-gptj-rm-25k"
More Information needed
corpus_1ACORPUS 1A: A corpus of President Clinton's terrorism-related discourse, for a historical research project.
Period: January 20, 1997 (Clinton’s second term inauguration day) – January 20, 2001 (the day Clinton left office).
Search parameters: All documents on the American Presidency Project site within the above timeframe returned through a keyword search ‘terror*’
using the wildcard star to return all variants, terrorism, terrorist, etc. Results were further refined to only those associated… See the full description on the dataset page: https://huggingface.co/datasets/GPT-JF/corpus_1A.Corpus_1BCORPUS 1B: A corpus of President George W. Bush's terrorism-related discourse, for a historical research project.
Period: September 11, 2001 (preferable to starting at the beginning of the Bush presidency, for a clean break between pre-9/11 and post-9/11.
-January 20, 2005 (the end of Bush’s first term).
Search parameters: All documents on the American Presidency Project site within the above timeframe returned through a keyword search ‘terror*’
using the wildcard star to return all variants… See the full description on the dataset page: https://huggingface.co/datasets/GPT-JF/Corpus_1B.gpt_judge_prefdatagpt_just_10gptjodoo4000GPT_judgement_sft_dpoGPT_judgement_base_vs_sftGPT_judgement_base_vs_dpo
