CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jerredchen00 /image-as-an-imu-finetuning Image as an IMU: Real-world Finetuning Dataset Official real-world finetuning dataset from Image as an IMU: Estimating Camera Motion from a Single Motion-Blurred Image (ICCV 2025 Oral). [arXiv] [Webpage] [GitHub] PIXL, University of Oxford Jerred Chen, Ronald Clark Dataset Details This dataset consists of 32 sequences of real-world motion-blurred videos in various indoor scenes, captured using the iPhone 13 camera. dataset_train_real-world.csv and… See the full description on the dataset page: https://huggingface.co/datasets/jerredchen00/image-as-an-imu-finetuning.image10K<n<100K0 likes4k downloads10mo agoHugging Face02rbhatia46 /embedding-finetuning-financeThis dataset can be used for fine-tuning embedding models using positive text pairs (question, context). text1K<n<10K5 likes260 downloads2y agoHugging Face03science-of-finetuning /diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04 Contains maximum activating examples for all the features of our crosscoder trained on gemma 2 2B layer 13 available here: https://huggingface.co/Butanium/gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04/blob/main/README.md base_examples.pt contains all the maximum examples of the feature on a subset of validation test of fineweb chat_examples.pt is the same but for lmsys chat data chat_base_examples.pt is a merge of the two above files. All files are of the type dict[int, list[tuple[float… See the full description on the dataset page: https://huggingface.co/datasets/science-of-finetuning/diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04.tabular10K<n<100K0 likes235 downloads1y agoHugging Face04soniawmeyer /travel-conversations-finetuning UltraChat Dataset (HuggingFace) For prototyping and model training, the project utilized the "UltraChat" dataset available from HuggingFace. This dataset comprises 10 JSONLines files, totaling 1.5 million conversations, each stored as lists of strings. The initial preprocessing involved standardizing the text data by converting it to lowercase, removing punctuation using regular expressions, and applying lemmatization with part-of-speech tagging. These steps ensured uniformity and… See the full description on the dataset page: https://huggingface.co/datasets/soniawmeyer/travel-conversations-finetuning.text10K<n<100K6 likes75 downloads2y agoHugging Face05soniawmeyer /reddit-travel-QA-finetuningThis dataset was sourced through a series of daily requests to the Reddit API, aiming to capture diverse and real-time travel-related discussions from multiple travel-related subreddits, sourced from this list: https://www.reddit.com/r/travel/comments/1100hca/the_definitive_list_of_travel_subreddits_to_help/, along with subreddits for common travel destinations. Requested was top 100 of the year, this was executed only one, then hot 50 daily. Data aggregation involved concatenating and… See the full description on the dataset page: https://huggingface.co/datasets/soniawmeyer/reddit-travel-QA-finetuning.tabular10K<n<100K5 likes58 downloads2y agoHugging Face06science-of-finetuning /diffing-stats-SAE-difference_cb-gemma-2-2b-L13-k100-x8-lr1e-04-local-shufflingtabular10K<n<100K0 likes54 downloads1y agoHugging Face07cheekymachine /enron_labeled_emails_with_subjects-llama2-7b_finetuningtexttext-classification1K<n<10K5 likes44 downloads3y agoHugging Face08open-paws /conversational-finetuning-llama-format Open Paws Conversational Finetuning Llama Format This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Training Data Format: CSV (Comma-separated values) Languages: Multilingual (primarily English) Focus: Animal advocacy and ethical reasoning Organization: Open Paws… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/conversational-finetuning-llama-format.texttext-generation10K<n<100K2 likes39 downloads1y agoHugging Face09MahdiAbdoZahra /RAG_vs_FineTuning_Comparison_Persian_V2text1K<n<10K1 likes36 downloads2d agoHugging Face10MahdiAbdoZahra /RAG_vs_FineTuning_Comparison_Persian_V1tabularn<1K1 likes29 downloads2d agoHugging Face11nulltella /bbc-articles-finetuning-classiftext1K<n<10K0 likes25 downloads3y agoHugging Face12science-of-finetuning /diffing-stats-Meta-Llama-3.1-8B-L16-mu2.0e-02-lr1e-04-local-shuffling-CCLosstabular100K<n<1M0 likes25 downloads1y agoHugging Face13science-of-finetuning /max-activating-examples-gemma-2-2b-l13-ckissanetabular10K<n<100K0 likes24 downloads2y agoHugging Face14williamkgao /FinetuningMoondream2tabularn<1K0 likes23 downloads2y agoHugging Face15cowWhySo /selenium-finetuning-datasetThis was created to finetune Gemini to take HTML and return a JSON with selenium selectors for extraction. The HTML was generated randomly using LLM and passed into LLM to generate selectors. All webpages are made up and don't exist. textn<1K0 likes23 downloads2y agoHugging Face16science-of-finetuning /diffing-stats-SAE-base-gemma-2-2b-L13-k100-x32-lr1e-04-local-shufflingtabular100K<n<1M0 likes22 downloads1y agoHugging Face17Dipe00 /Urgency-tone-topic-on-enron_labeled_emails_with_subjects-llama2-7b_finetuningtext1K<n<10K0 likes20 downloads2y agoHugging Face18falan42 /o1-medical-finetuning-trtext1K<n<10K1 likes20 downloads1y agoHugging Face19science-of-finetuning /diffing-stats-gemma-2-2b-L13-k100-lr1e-04-local-shuffling-CCLosstabular10K<n<100K0 likes19 downloads1y agoHugging Face20science-of-finetuning /diffing-stats-SAE-difference_cb-gemma-2-2b-L13-k100-lr1e-04-local-shufflingtabular10K<n<100K0 likes19 downloads1y agoHugging Face21science-of-finetuning /diffing-stats-SAE-difference_cb-gemma-2-2b-L13-k100-x2-lr1e-04-local-shufflingtabular1K<n<10K0 likes19 downloads1y agoHugging Face22science-of-finetuning /diffing-stats-SAE-chat-gemma-2-2b-L13-k100-lr1e-04-local-shufflingtabular10K<n<100K0 likes18 downloads1y agoHugging Face23science-of-finetuning /diffing-stats-qwen3_1_7B-kansas_abortion-L14-Crosscoder-s2-t100-k100-lr1e-04-x32tabular10K<n<100K0 likes17 downloads1y agoHugging Face24SohamNale /Banking_Dataset_for_LLM_Finetuningtext1K<n<10K1 likes16 downloads3y agoHugging Face25science-of-finetuning /diffing-stats-SAE-difference-gemma-2-2b-L13-k100-lr1e-04-local-shufflingtabular10K<n<100K0 likes16 downloads1y agoHugging Face26MasterControlAIML /Medmcqa-For-FinetuningQwentext100K<n<1M0 likes15 downloads2y agoHugging Face27thudoann /finetuningllmtextquestion-answering10K<n<100K0 likes14 downloads3y agoHugging Face28Essacheez /Knowledge_blanks_finetuningtextn<1K0 likes14 downloads1y agoHugging Face29science-of-finetuning /diffing-stats-SAEdiff_ftb-qwen3_1_7B-kansas_abortion-L14-s1-t200-k100-lr1e-04-x32tabular10K<n<100K0 likes14 downloads1y agoHugging Face30science-of-finetuning /diffing-stats-qwen3_1_7B-comment_cake_bake-L14-Crosscoder-s2-t100-k100-lr1e-04-x32tabular10K<n<100K0 likes14 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.