datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpt2-training-ar-zh-ko-ja-4b
Balanced Arabic-Chinese-Korean-Japanese 4B-token training data
Sequential FineWeb-2 documents tokenized with CausalNLP/gpt2-tokenizer-ar-zh-ko-ja. Each language contains at least 1,000,000,000 tokens in complete documents. Total target: 4,000,000,000 tokens. Approximately 100,000,000 tokens per Parquet shard.
Splits: arb_Arab, cmn_Hani, kor_Hang, jpn_Jpan. Schema: text: string, input_ids: list<int32>.
araproje_hellaswag_tr_conf_gpt2_bestscore_is
Dataset Card for "araproje_hellaswag_tr_conf_gpt2_bestscore_is"
More Information needed
GPT2_TEST-train-datasetaraproje_hellaswag_tr_conf_gpt2_nearestscore_true_y
Dataset Card for "araproje_hellaswag_tr_conf_gpt2_nearestscore_true_y"
More Information needed
araproje_hellaswag_tr_conf_gpt2_worstscore_reversed
Dataset Card for "araproje_hellaswag_tr_conf_gpt2_worstscore_reversed"
More Information needed
araproje_hellaswag_tr_conf_gpt2_bestscore
Dataset Card for "araproje_hellaswag_tr_conf_gpt2_bestscore"
More Information needed
araproje_arc_tr_conf_gpt2_nearestscore_true
Dataset Card for "araproje_arc_tr_conf_gpt2_nearestscore_true"
More Information needed
FINGPT_QA_GPT2_FP16-train-datasetaraproje_hellaswag_tr_conf_gpt2_worstscore
Dataset Card for "araproje_hellaswag_tr_conf_gpt2_worstscore"
More Information needed
araproje_arc_tr_conf_gpt2_nearestscore_true_y
Dataset Card for "araproje_arc_tr_conf_gpt2_nearestscore_true_y"
More Information needed
dstc2_dialogues_transcript_input_gpt2multi_nli_negation_duplication_gpt2-tokenized_trainaraproje_hellaswag_tr_conf_gpt2_bestscore_reversed
Dataset Card for "araproje_hellaswag_tr_conf_gpt2_bestscore_reversed"
More Information needed
araproje_mmlu_tr_conf_gpt2_nearestscore_true_x
Dataset Card for "araproje_mmlu_tr_conf_gpt2_nearestscore_true_x"
More Information needed
araproje_arc_tr_conf_gpt2_nearestscore_true_x
Dataset Card for "araproje_arc_tr_conf_gpt2_nearestscore_true_x"
More Information needed
araproje_hellaswag_tr_conf_gpt2_nearestscore_true_x
Dataset Card for "araproje_hellaswag_tr_conf_gpt2_nearestscore_true_x"
More Information needed
c4_val_tinyllama_gpt2-medium_tr0.8_val1000araproje_mmlu_tr_conf_gpt2_nearestscore_true
Dataset Card for "araproje_mmlu_tr_conf_gpt2_nearestscore_true"
More Information needed
araproje_hellaswag_tr_conf_gpt2_nearestscore_true
Dataset Card for "araproje_hellaswag_tr_conf_gpt2_nearestscore_true"
More Information needed
gpt2-qa-train-dsaraproje_mmlu_tr_conf_gpt2_nearestscore_true_y
Dataset Card for "araproje_mmlu_tr_conf_gpt2_nearestscore_true_y"
More Information needed
lambada-openai-train-tokenized-gpt2gpt2traindata
