CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Korakoe /Vicuna-Uncleaned-Alpaca-Formattexttext-generation10K<n<100K4 likes54 downloads3y agoHugging Face02Vezora /Mini_Orca_Code_Uncencored_alpaca_FormatThis is dataset is a modified version of "psmathur's" Mini orca dataset, formated in the alpaca format and uncencored. This dataset is filtered to only feature coding instructions around 50k code examples. For ALPACA LORA users: Modules you can target with lora:"gate_proj", "down_proj", "up_proj", "q_proj", "v_proj", "k_proj", "o_proj" Most lora models use:"q_proj", "v_proj", "k_proj", "o_proj" Platypus which got terrific results: "gate_proj", "down_proj", "up_proj" Research on targeting… See the full description on the dataset page: https://huggingface.co/datasets/Vezora/Mini_Orca_Code_Uncencored_alpaca_Format.text10K<n<100K1 likes22 downloads3y agoHugging Face03adambuttrick /100K_deduplicated_ner_indexes_name_country_alpaca_format_json_response_all_casestext100K<n<1M0 likes22 downloads3y agoHugging Face04Vezora /Dolphin1m_gpt4_Alpaca_formattext100K<n<1M1 likes21 downloads3y agoHugging Face05Fredithefish /ShareGPT-unfiltered-alpaca-lora-formattext10K<n<100K3 likes18 downloads3y agoHugging Face06Vezora /Gorilla_Alpaca_FormatThis is the dataset to the model used to train gorilla 7b but in the alpaca format, for lora training. Thank you to microsoft and uc berkly for open sourcing these datasets. As of now I do not believe this dataset works, will have to do more testing, but gorilla team plans to realease training code which might make it easer to see how this was fully done. and how it can be done with lora. For ALPACA LORA users: Modules you can target with lora:"gate_proj", "down_proj", "up_proj", "q_proj"… See the full description on the dataset page: https://huggingface.co/datasets/Vezora/Gorilla_Alpaca_Format.text1K<n<10K0 likes17 downloads3y agoHugging Face07ysn-rfd /fibonacci_alpaca_to_sharegpt_gpt_format_convert_new_dataset_releasetext1K<n<10K3 likes15 downloads1y agoHugging Face08adambuttrick /50K_deduplicated_ner_indexes_name_country_alpaca_format_json_responsetext10K<n<100K1 likes14 downloads3y agoHugging Face09Ruoyao /gsm8k_reasoning_paths_deepseek_alpaca_formattext1K<n<10K0 likes13 downloads1y agoHugging Face10Nekochu /novel17_train_alpaca_formatCredit: AlexanderDoria/novel17_test text1K<n<10K2 likes12 downloads3y agoHugging Face11adambuttrick /100K-ner-indexes-multiple-organizations-locations-alpaca-format-json-response-all-casestext100K<n<1M1 likes12 downloads3y agoHugging Face12kdunee /IntentGuard-2-alpaca-format IntentGuard 2 Alpaca Format This dataset is IntentGuard-2 converted to Alpaca instruction format. Every row uses instruction: "/intentguard". Source Files IntentGuard-2/train.jsonl IntentGuard-2/test.jsonl Output Files train.json test.json Both output files are JSON arrays. Each item contains instruction, input, and output. Rebuild Run python3 scripts/convert_to_alpaca.py from this directory. texttext-generation1K<n<10K0 likes12 downloads3mo agoHugging Face13ysn-rfd /Persian_Text_Dataset_QA_Alpaca_Format_Conversations5text100K<n<1M1 likes11 downloads1y agoHugging Face14Laurie /HC3-Chinese-AlpacaFormattext1K<n<10K1 likes10 downloads3y agoHugging Face15adambuttrick /500K-ner-indexes-multiple-organizations-locations-alpaca-format-json-response-all-casestext100K<n<1M1 likes9 downloads3y agoHugging Face16artemkoverchik /taxonomies-dataset-alpaca-prompt-formattextn<1K0 likes9 downloads2y agoHugging Face17Ruoyao /gsm8k_reasoning_paths_test_deepseek_alpaca_formattext1K<n<10K0 likes9 downloads1y agoHugging Face18ARFai /GSM8K_ALPACA_FORMATtext10K<n<100K0 likes8 downloads2y agoHugging Face19EmTpro01 /ingredients-recipe-alpaca-formattext100K<n<1M0 likes7 downloads2y agoHugging Face20couchpotato888 /dolly-and-alpaca-lora-data-formattext10K<n<100K1 likes6 downloads3y agoHugging Face21ysn-rfd /fibonacci_alpaca_to_gemma_format_dataset_Fibonacci_ai_persian_pythontext1K<n<10K0 likes6 downloads1y agoHugging Face22adambuttrick /50K_ner_indexes_name_country_alpaca_formattext10K<n<100K0 likes5 downloads3y agoHugging Face23ChrisSacrumCor /alpacaformatGKChestertontextn<1K0 likes5 downloads2y agoHugging Face24dasarath /EZYBill_SQL_alpaca_formattextn<1K0 likes4 downloads2y agoHugging Face25kdunee /IntentGuard-1-alpaca-formatThis is the IntentGuard-1 dataset converted to alpaca instruction format. texttext-generation1K<n<10K0 likes4 downloads2y agoHugging Face26usham /mental-alpaca-formattext100K<n<1M0 likes4 downloads1y agoHugging Face27abdullahyasir /Alpaca_format_datasettextn<1K0 likes3 downloads2y agoHugging Face28Vezora /news_seniment_gpt_alpacaformatThis dataset is a alpaca formatted version of "oliverwang15/news_with_gpt_instructions" (https://huggingface.co/datasets/oliverwang15/news_with_gpt_instructions) 20k examples of grading senitment using gpt (unclear which model) (used to train fingptv3). For ALPACA LORA users: Modules you can target with lora:"gate_proj", "down_proj", "up_proj", "q_proj", "v_proj", "k_proj", "o_proj" Most lora models use:"q_proj", "v_proj", "k_proj", "o_proj" Platypus which got terrific results: "gate_proj"… See the full description on the dataset page: https://huggingface.co/datasets/Vezora/news_seniment_gpt_alpacaformat.text10K<n<100K0 likes2 downloads3y agoHugging Face29mohitkunder /autorater-alpaca-format-50ktext10K<n<100K0 likes2 downloads2y agoHugging Face30AndreiMuresanu /alpaca_flan-formattext10K<n<100K2 likes1 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.