datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
image-as-an-imu-finetuning
Image as an IMU: Real-world Finetuning Dataset
Official real-world finetuning dataset from Image as an IMU: Estimating Camera Motion from a Single Motion-Blurred Image (ICCV 2025 Oral).
[arXiv] [Webpage] [GitHub]
PIXL, University of Oxford
Jerred Chen, Ronald Clark
Dataset Details
This dataset consists of 32 sequences of real-world motion-blurred videos in various indoor scenes, captured using the iPhone 13 camera.
dataset_train_real-world.csv and… See the full description on the dataset page: https://huggingface.co/datasets/jerredchen00/image-as-an-imu-finetuning.embedding-finetuning-financeThis dataset can be used for fine-tuning embedding models using positive text pairs (question, context).
diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04 Contains maximum activating examples for all the features of our crosscoder trained on gemma 2 2B layer 13 available here: https://huggingface.co/Butanium/gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04/blob/main/README.md
base_examples.pt contains all the maximum examples of the feature on a subset of validation test of fineweb
chat_examples.pt is the same but for lmsys chat data
chat_base_examples.pt is a merge of the two above files.
All files are of the type dict[int, list[tuple[float… See the full description on the dataset page: https://huggingface.co/datasets/science-of-finetuning/diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04.travel-conversations-finetuning
UltraChat Dataset (HuggingFace)
For prototyping and model training, the project utilized the "UltraChat" dataset available from HuggingFace. This dataset comprises 10 JSONLines files, totaling 1.5 million conversations, each stored as lists of strings. The initial preprocessing involved standardizing the text data by converting it to lowercase, removing punctuation using regular expressions, and applying lemmatization with part-of-speech tagging. These steps ensured uniformity and… See the full description on the dataset page: https://huggingface.co/datasets/soniawmeyer/travel-conversations-finetuning.reddit-travel-QA-finetuningThis dataset was sourced through a series of daily requests to the Reddit API, aiming to capture diverse and real-time travel-related discussions from multiple travel-related subreddits, sourced from this list: https://www.reddit.com/r/travel/comments/1100hca/the_definitive_list_of_travel_subreddits_to_help/, along with subreddits for common travel destinations. Requested was top 100 of the year, this was executed only one, then hot 50 daily. Data aggregation involved concatenating and… See the full description on the dataset page: https://huggingface.co/datasets/soniawmeyer/reddit-travel-QA-finetuning.diffing-stats-SAE-difference_cb-gemma-2-2b-L13-k100-x8-lr1e-04-local-shufflingenron_labeled_emails_with_subjects-llama2-7b_finetuningconversational-finetuning-llama-format
Open Paws Conversational Finetuning Llama Format
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Training Data
Format: CSV (Comma-separated values)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning
Organization: Open Paws… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/conversational-finetuning-llama-format.RAG_vs_FineTuning_Comparison_Persian_V2RAG_vs_FineTuning_Comparison_Persian_V1bbc-articles-finetuning-classifdiffing-stats-Meta-Llama-3.1-8B-L16-mu2.0e-02-lr1e-04-local-shuffling-CCLossmax-activating-examples-gemma-2-2b-l13-ckissaneFinetuningMoondream2selenium-finetuning-datasetThis was created to finetune Gemini to take HTML and return a JSON with selenium selectors for extraction.
The HTML was generated randomly using LLM and passed into LLM to generate selectors. All webpages are made up and don't exist.
diffing-stats-SAE-base-gemma-2-2b-L13-k100-x32-lr1e-04-local-shufflingUrgency-tone-topic-on-enron_labeled_emails_with_subjects-llama2-7b_finetuningo1-medical-finetuning-trdiffing-stats-gemma-2-2b-L13-k100-lr1e-04-local-shuffling-CCLossdiffing-stats-SAE-difference_cb-gemma-2-2b-L13-k100-lr1e-04-local-shufflingdiffing-stats-SAE-difference_cb-gemma-2-2b-L13-k100-x2-lr1e-04-local-shufflingdiffing-stats-SAE-chat-gemma-2-2b-L13-k100-lr1e-04-local-shufflingdiffing-stats-qwen3_1_7B-kansas_abortion-L14-Crosscoder-s2-t100-k100-lr1e-04-x32Banking_Dataset_for_LLM_Finetuningdiffing-stats-SAE-difference-gemma-2-2b-L13-k100-lr1e-04-local-shufflingMedmcqa-For-FinetuningQwenfinetuningllmKnowledge_blanks_finetuningdiffing-stats-SAEdiff_ftb-qwen3_1_7B-kansas_abortion-L14-s1-t200-k100-lr1e-04-x32diffing-stats-qwen3_1_7B-comment_cake_bake-L14-Crosscoder-s2-t100-k100-lr1e-04-x32
