CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jerredchen00 /image-as-an-imu-finetuning Image as an IMU: Real-world Finetuning Dataset Official real-world finetuning dataset from Image as an IMU: Estimating Camera Motion from a Single Motion-Blurred Image (ICCV 2025 Oral). [arXiv] [Webpage] [GitHub] PIXL, University of Oxford Jerred Chen, Ronald Clark Dataset Details This dataset consists of 32 sequences of real-world motion-blurred videos in various indoor scenes, captured using the iPhone 13 camera. dataset_train_real-world.csv and… See the full description on the dataset page: https://huggingface.co/datasets/jerredchen00/image-as-an-imu-finetuning.image10K<n<100K0 likes3.8k downloads10mo agoHugging Face02rbhatia46 /embedding-finetuning-financeThis dataset can be used for fine-tuning embedding models using positive text pairs (question, context). text1K<n<10K5 likes259 downloads2y agoHugging Face03science-of-finetuning /diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04 Contains maximum activating examples for all the features of our crosscoder trained on gemma 2 2B layer 13 available here: https://huggingface.co/Butanium/gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04/blob/main/README.md base_examples.pt contains all the maximum examples of the feature on a subset of validation test of fineweb chat_examples.pt is the same but for lmsys chat data chat_base_examples.pt is a merge of the two above files. All files are of the type dict[int, list[tuple[float… See the full description on the dataset page: https://huggingface.co/datasets/science-of-finetuning/diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04.tabular10K<n<100K0 likes238 downloads1y agoHugging Face04soniawmeyer /travel-conversations-finetuning UltraChat Dataset (HuggingFace) For prototyping and model training, the project utilized the "UltraChat" dataset available from HuggingFace. This dataset comprises 10 JSONLines files, totaling 1.5 million conversations, each stored as lists of strings. The initial preprocessing involved standardizing the text data by converting it to lowercase, removing punctuation using regular expressions, and applying lemmatization with part-of-speech tagging. These steps ensured uniformity and… See the full description on the dataset page: https://huggingface.co/datasets/soniawmeyer/travel-conversations-finetuning.text10K<n<100K6 likes73 downloads2y agoHugging Face05Mreeb /Dermatology-Question-Answer-Dataset-For-Fine-Tuning Dataset Details The data set has about 1 Million Tokens for Training and about 1500 question answers. Dataset Description This dataset is a comprehensive compilation of questions related to dermatology, spanning inquiries about various skin diseases, their symptoms, recommended medications, and available treatment modalities. Each question is paired with a concise and informative response, making it an ideal resource for training and fine-tuning language models in the… See the full description on the dataset page: https://huggingface.co/datasets/Mreeb/Dermatology-Question-Answer-Dataset-For-Fine-Tuning.tabulartext-generation1K<n<10K7 likes69 downloads3y agoHugging Face06MahdiAbdoZahra /RAG_vs_FineTuning_Comparison_Persian_V2text1K<n<10K1 likes50 downloads5d agoHugging Face07soniawmeyer /reddit-travel-QA-finetuningThis dataset was sourced through a series of daily requests to the Reddit API, aiming to capture diverse and real-time travel-related discussions from multiple travel-related subreddits, sourced from this list: https://www.reddit.com/r/travel/comments/1100hca/the_definitive_list_of_travel_subreddits_to_help/, along with subreddits for common travel destinations. Requested was top 100 of the year, this was executed only one, then hot 50 daily. Data aggregation involved concatenating and… See the full description on the dataset page: https://huggingface.co/datasets/soniawmeyer/reddit-travel-QA-finetuning.tabular10K<n<100K5 likes48 downloads2y agoHugging Face08cheekymachine /enron_labeled_emails_with_subjects-llama2-7b_finetuningtexttext-classification1K<n<10K5 likes44 downloads3y agoHugging Face09MahdiAbdoZahra /RAG_vs_FineTuning_Comparison_Persian_V1tabularn<1K1 likes43 downloads5d agoHugging Face10open-paws /conversational-finetuning-llama-format Open Paws Conversational Finetuning Llama Format This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Training Data Format: CSV (Comma-separated values) Languages: Multilingual (primarily English) Focus: Animal advocacy and ethical reasoning Organization: Open Paws… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/conversational-finetuning-llama-format.texttext-generation10K<n<100K2 likes41 downloads1y agoHugging Face11Dipe00 /Urgency-tone-topic-on-enron_labeled_emails_with_subjects-llama2-7b_finetuningtext1K<n<10K0 likes29 downloads2y agoHugging Face12williamkgao /FinetuningMoondream2tabularn<1K0 likes27 downloads2y agoHugging Face13nulltella /bbc-articles-finetuning-classiftext1K<n<10K0 likes25 downloads3y agoHugging Face14science-of-finetuning /diffing-stats-Meta-Llama-3.1-8B-L16-mu2.0e-02-lr1e-04-local-shuffling-CCLosstabular100K<n<1M0 likes25 downloads1y agoHugging Face15science-of-finetuning /max-activating-examples-gemma-2-2b-l13-ckissanetabular10K<n<100K0 likes23 downloads2y agoHugging Face16cowWhySo /selenium-finetuning-datasetThis was created to finetune Gemini to take HTML and return a JSON with selenium selectors for extraction. The HTML was generated randomly using LLM and passed into LLM to generate selectors. All webpages are made up and don't exist. textn<1K0 likes22 downloads2y agoHugging Face17falan42 /o1-medical-finetuning-trtext1K<n<10K1 likes18 downloads1y agoHugging Face18SohamNale /Banking_Dataset_for_LLM_Finetuningtext1K<n<10K1 likes16 downloads3y agoHugging Face19MasterControlAIML /Medmcqa-For-FinetuningQwentext100K<n<1M0 likes15 downloads2y agoHugging Face20science-of-finetuning /diffing-stats-gemma-2-2b-L13-k100-lr1e-04-local-shuffling-CCLosstabular10K<n<100K0 likes15 downloads1y agoHugging Face21science-of-finetuning /diffing-stats-qwen3_1_7B-kansas_abortion-L14-Crosscoder-s2-t100-k100-lr1e-04-x32tabular10K<n<100K0 likes15 downloads1y agoHugging Face22Essacheez /Knowledge_blanks_finetuningtextn<1K0 likes14 downloads1y agoHugging Face23thudoann /finetuningllmtextquestion-answering10K<n<100K0 likes11 downloads3y agoHugging Face24thudoann /finetuningllm2texttable-question-answering10K<n<100K0 likes11 downloads3y agoHugging Face25srirammoorthi /finetuningdatatabular10K<n<100K0 likes11 downloads2y agoHugging Face26lorixmassello /Akka_Finetuning_Llama3.2textquestion-answeringn<1K0 likes11 downloads2y agoHugging Face27ThrishaSivasakthi /Tamil-Finetuning-data Dataset Card for Dataset Name This dataset is designed for fine-tuning Large Language Models (LLMs) in Tamil, enabling them to understand and generate high-quality Tamil text across multiple domains. It contains 72,000 curated and generated samples, ensuring a rich linguistic diversity that improves model generalization. 🔹 Sources: Kaggle Tamil NLP, Sentiment Analysis datasets, and synthetic data. 🔹 Languages: Tamil, Tanglish (Tamil-English mix), and regional Tamil dialects. 🔹… See the full description on the dataset page: https://huggingface.co/datasets/ThrishaSivasakthi/Tamil-Finetuning-data.text10K<n<100K0 likes11 downloads2y agoHugging Face28science-of-finetuning /diffing-stats-gemma3_1B-kansas_abortion-L19-k100-lr1e-03-x32-local-shuffling-Crosscodertabular10K<n<100K0 likes11 downloads1y agoHugging Face29jr303 /detection_sexist_text_DPO_fine-tuning_formeted_datasettext1K<n<10K0 likes10 downloads2y agoHugging Face30drgary /uslaw_for_finetuningtextn<1K3 likes10 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.