CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Korakoe /Vicuna-Uncleaned-Alpaca-Formattexttext-generation10K<n<100K4 likes53 downloads3y agoHugging Face02adambuttrick /100K_deduplicated_ner_indexes_name_country_alpaca_format_json_response_all_casestext100K<n<1M0 likes26 downloads3y agoHugging Face03Vezora /Dolphin1m_gpt4_Alpaca_formattext100K<n<1M1 likes23 downloads3y agoHugging Face04Vezora /Gorilla_Alpaca_FormatThis is the dataset to the model used to train gorilla 7b but in the alpaca format, for lora training. Thank you to microsoft and uc berkly for open sourcing these datasets. As of now I do not believe this dataset works, will have to do more testing, but gorilla team plans to realease training code which might make it easer to see how this was fully done. and how it can be done with lora. For ALPACA LORA users: Modules you can target with lora:"gate_proj", "down_proj", "up_proj", "q_proj"… See the full description on the dataset page: https://huggingface.co/datasets/Vezora/Gorilla_Alpaca_Format.text1K<n<10K0 likes21 downloads3y agoHugging Face05Vezora /Mini_Orca_Code_Uncencored_alpaca_FormatThis is dataset is a modified version of "psmathur's" Mini orca dataset, formated in the alpaca format and uncencored. This dataset is filtered to only feature coding instructions around 50k code examples. For ALPACA LORA users: Modules you can target with lora:"gate_proj", "down_proj", "up_proj", "q_proj", "v_proj", "k_proj", "o_proj" Most lora models use:"q_proj", "v_proj", "k_proj", "o_proj" Platypus which got terrific results: "gate_proj", "down_proj", "up_proj" Research on targeting… See the full description on the dataset page: https://huggingface.co/datasets/Vezora/Mini_Orca_Code_Uncencored_alpaca_Format.text10K<n<100K1 likes19 downloads3y agoHugging Face06Fredithefish /ShareGPT-unfiltered-alpaca-lora-formattext10K<n<100K3 likes18 downloads3y agoHugging Face07Ruoyao /gsm8k_reasoning_paths_deepseek_alpaca_formattext1K<n<10K0 likes17 downloads1y agoHugging Face08kdunee /IntentGuard-2-alpaca-format IntentGuard 2 Alpaca Format This dataset is IntentGuard-2 converted to Alpaca instruction format. Every row uses instruction: "/intentguard". Source Files IntentGuard-2/train.jsonl IntentGuard-2/test.jsonl Output Files train.json test.json Both output files are JSON arrays. Each item contains instruction, input, and output. Rebuild Run python3 scripts/convert_to_alpaca.py from this directory. texttext-generation1K<n<10K0 likes16 downloads3mo agoHugging Face09ysn-rfd /fibonacci_alpaca_to_sharegpt_gpt_format_convert_new_dataset_releasetext1K<n<10K3 likes15 downloads1y agoHugging Face10adambuttrick /50K_deduplicated_ner_indexes_name_country_alpaca_format_json_responsetext10K<n<100K1 likes14 downloads3y agoHugging Face11Nekochu /novel17_train_alpaca_formatCredit: AlexanderDoria/novel17_test text1K<n<10K2 likes12 downloads3y agoHugging Face12adambuttrick /100K-ner-indexes-multiple-organizations-locations-alpaca-format-json-response-all-casestext100K<n<1M1 likes12 downloads3y agoHugging Face13Ruoyao /gsm8k_reasoning_paths_test_deepseek_alpaca_formattext1K<n<10K0 likes11 downloads1y agoHugging Face14ysn-rfd /Persian_Text_Dataset_QA_Alpaca_Format_Conversations5text100K<n<1M1 likes11 downloads11mo agoHugging Face15Laurie /HC3-Chinese-AlpacaFormattext1K<n<10K1 likes10 downloads3y agoHugging Face16ARFai /GSM8K_ALPACA_FORMATtext10K<n<100K0 likes10 downloads2y agoHugging Face17adambuttrick /500K-ner-indexes-multiple-organizations-locations-alpaca-format-json-response-all-casestext100K<n<1M1 likes9 downloads3y agoHugging Face18artemkoverchik /taxonomies-dataset-alpaca-prompt-formattextn<1K0 likes9 downloads2y agoHugging Face19EmTpro01 /ingredients-recipe-alpaca-formattext100K<n<1M0 likes7 downloads2y agoHugging Face20couchpotato888 /dolly-and-alpaca-lora-data-formattext10K<n<100K1 likes6 downloads3y agoHugging Face21ysn-rfd /fibonacci_alpaca_to_gemma_format_dataset_Fibonacci_ai_persian_pythontext1K<n<10K0 likes6 downloads1y agoHugging Face22adambuttrick /50K_ner_indexes_name_country_alpaca_formattext10K<n<100K0 likes5 downloads3y agoHugging Face23ChrisSacrumCor /alpacaformatGKChestertontextn<1K0 likes5 downloads2y agoHugging Face24dasarath /EZYBill_SQL_alpaca_formattextn<1K0 likes4 downloads2y agoHugging Face25kdunee /IntentGuard-1-alpaca-formatThis is the IntentGuard-1 dataset converted to alpaca instruction format. texttext-generation1K<n<10K0 likes4 downloads2y agoHugging Face26usham /mental-alpaca-formattext100K<n<1M0 likes4 downloads1y agoHugging Face27abdullahyasir /Alpaca_format_datasettextn<1K0 likes3 downloads2y agoHugging Face28Vezora /news_seniment_gpt_alpacaformatThis dataset is a alpaca formatted version of "oliverwang15/news_with_gpt_instructions" (https://huggingface.co/datasets/oliverwang15/news_with_gpt_instructions) 20k examples of grading senitment using gpt (unclear which model) (used to train fingptv3). For ALPACA LORA users: Modules you can target with lora:"gate_proj", "down_proj", "up_proj", "q_proj", "v_proj", "k_proj", "o_proj" Most lora models use:"q_proj", "v_proj", "k_proj", "o_proj" Platypus which got terrific results: "gate_proj"… See the full description on the dataset page: https://huggingface.co/datasets/Vezora/news_seniment_gpt_alpacaformat.text10K<n<100K0 likes2 downloads3y agoHugging Face29mohitkunder /autorater-alpaca-format-50ktext10K<n<100K0 likes2 downloads2y agoHugging Face30AndreiMuresanu /alpaca_flan-formattext10K<n<100K2 likes1 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.