CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hatemestinbejaia /ExperimentDATA_knowledge_distillation_vs_fine_tuningtabular100M<n<1B1 likes64k downloads9mo agoHugging Face02jerredchen00 /image-as-an-imu-finetuning Image as an IMU: Real-world Finetuning Dataset Official real-world finetuning dataset from Image as an IMU: Estimating Camera Motion from a Single Motion-Blurred Image (ICCV 2025 Oral). [arXiv] [Webpage] [GitHub] PIXL, University of Oxford Jerred Chen, Ronald Clark Dataset Details This dataset consists of 32 sequences of real-world motion-blurred videos in various indoor scenes, captured using the iPhone 13 camera. dataset_train_real-world.csv and… See the full description on the dataset page: https://huggingface.co/datasets/jerredchen00/image-as-an-imu-finetuning.image10K<n<100K0 likes3.9k downloads10mo agoHugging Face03RicemanT /Anime-Background-Finetuning-V1.1 Anime-Background-Finetuning (10143 manually curated by hand images from danbooru and reddit collections) The dataset contain roughly 2k of anime Screencap data and 8k of scrapped danbooru illustration data. This is the proccessed version of the dataset meant to be used for my personal finetuning practice project, please visit my RicemanT/Background-Finetuning repo for the raw unprocessed data that you can process yourself. The dataset have two minor type of processing being done… See the full description on the dataset page: https://huggingface.co/datasets/RicemanT/Anime-Background-Finetuning-V1.1.image10K<n<100K7 likes3.3k downloads3mo agoHugging Face04HappyHenAi /Anime-Background-Finetuning-V1.1 Anime-Background-Finetuning (10143 manually curated by hand images from danbooru and reddit collections) The dataset contain roughly 2k of anime Screencap data and 8k of scrapped danbooru illustration data. This is the proccessed version of the dataset meant to be used for my personal finetuning practice project, please visit my RicemanT/Background-Finetuning repo for the raw unprocessed data that you can process yourself. The dataset have two minor type of processing being done… See the full description on the dataset page: https://huggingface.co/datasets/HappyHenAi/Anime-Background-Finetuning-V1.1.image10K<n<100K4 likes2.7k downloads1mo agoHugging Face05lightonai /embeddings-fine-tuning Overview This dataset is composed of high quality data sources with mined hard negatives. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using this dataset or its curated version. This dataset has originally been created to follow the nv-retrieve setup, that mines the closest negatives to the query in a dataset and filter false negatives if their bi-encoder similarity is higher than a… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning.text10M<n<100M24 likes2.1k downloads3mo agoHugging Face06RicemanT /Anime-Background-Finetuning-Unprocessed Anime-Background-Dataset (10143 manually curated by hand images from danbooru and reddit collections) The dataset contain roughly 2k of Screencap data and 8k of scrapped danbooru illustration data. It is all raw unprocessed data, the illust folder contain scrapped danbooru tags sidecar .txt on most of the images, while the screencap have non. The processed data is being worked on a seperate repo (Anime-Background-Finetuning) 10K<n<100K3 likes2k downloads3mo agoHugging Face07sveneziale /finetuning-checkpointstext10K<n<100K0 likes1.7k downloads2d agoHugging Face08alkzar90 /ddpm-rl-finetuning-evals Dataset Card for Eval Finetuning Diffusion Models with Reinforcement Learning XYZ image10K<n<100K1 likes1.6k downloads2y agoHugging Face09Voxel51 /scanned-images-dataset-for-ocr-and-vlm-finetuning Dataset Card for scanned_images_dataset This is a FiftyOne dataset containing 3,482 scanned document images across 10 diverse document categories. Designed for OCR training and Vision-Language Model (VLM) fine-tuning, this dataset features real-world scanned documents with varied layouts, scanning quality, and document types. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/scanned-images-dataset-for-ocr-and-vlm-finetuning.imageimage-classification1K<n<10K2 likes1.5k downloads8mo agoHugging Face10lightonai /embeddings-fine-tuning-multilingual-unfiltered Overview This dataset provides multilingual and code retrieval data for fine-tuning text embedding models. It is composed of high quality data sources with mined documents annotated with bi-encoder scores. For each query, the 2048 closest documents are mined with snowflake-arctic-embed-l-v2.0 for MIRACL and MLDR and with gte-modernbert-base for CodeEditSearchTrain, and annotated with their bi-encoder similarity score. No false-negative filtering or cross-encoder annotation is… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-multilingual-unfiltered.4 likes1.4k downloads2mo agoHugging Face11Lycolys /fine-tuning-experiments-0820230 likes1.3k downloads3y agoHugging Face12huggingface-course /supervised-finetuning_quiz_student_responsestextn<1K4 likes1.2k downloads3h agoHugging Face13thetrillioniar /claude-sonnet-4.6-opus-4.8-mythos-5-fable-5-openai-finetuning-dataset OpenAI-Compatible Dataset Collection A collection of 29 datasets converted to OpenAI fine-tuning format ({"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}). Summary Metric Value Total Datasets 29 Total Rows ~1.5M Total Size ~1.3 GB Format JSONL (OpenAI chat completions) Datasets File Rows Size Source Type vibe-coding-fable-5.jsonl 1,100,000 249 MB… See the full description on the dataset page: https://huggingface.co/datasets/thetrillioniar/claude-sonnet-4.6-opus-4.8-mythos-5-fable-5-openai-finetuning-dataset.3 likes1.1k downloads3mo agoHugging Face14lightonai /embeddings-fine-tuning-filtered-en Overview This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets. The negatives were mined following the NV-Retriever setup: the closest documents to each query are mined as negatives, and false negatives are filtered out if… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-en.6 likes1.1k downloads2mo agoHugging Face15KaLM-Embedding /KaLM-embedding-finetuning-dataThe pretraining dataset is available at this link: HIT-TMG/KaLM-embedding-pretrain-data. Languages English, Chinese, Multilingual Dataset Structure Each in datasets is in the following format: query, string, one query per sample pos, list[string], usually containing one positive example neg, list[string], usually containing seven negative examples Dataset Summary All these datasets have been preprocessed and can be used for finetuning your embedding models.… See the full description on the dataset page: https://huggingface.co/datasets/KaLM-Embedding/KaLM-embedding-finetuning-data.textfeature-extraction1M<n<10M32 likes1.1k downloads10mo agoHugging Face16lightonai /embeddings-fine-tuning-filtered-ar Overview This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets. All splits except MIRACL and MLDR were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-ar.2 likes1k downloads2mo agoHugging Face17lightonai /embeddings-fine-tuning-filtered-fr Overview This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets. All splits except MIRACL and MLDR were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-fr.2 likes914 downloads2mo agoHugging Face18SimbaMaw1547 /south-african-finetuning South African Finetuning Datasets This dataset collection contains various NLP tasks for South African languages, organized by task and language. Dataset Structure The dataset follows this structure: task_name/ language_code/ train.jsonl dev.jsonl test.jsonl metadata.json Tasks This collection includes the following tasks: afrisent-semeval Language Train Validation Test tso 804 203 254… See the full description on the dataset page: https://huggingface.co/datasets/SimbaMaw1547/south-african-finetuning.0 likes874 downloads1y agoHugging Face19Johnblick187 /claude-sonnet-4.6-opus-4.8-mythos-5-fable-5-openai-finetuning-dataset OpenAI-Compatible Dataset Collection A collection of 29 datasets converted to OpenAI fine-tuning format ({"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}). Summary Metric Value Total Datasets 29 Total Rows ~1.5M Total Size ~1.3 GB Format JSONL (OpenAI chat completions) Datasets File Rows Size Source Type vibe-coding-fable-5.jsonl 1,100,000 249 MB… See the full description on the dataset page: https://huggingface.co/datasets/Johnblick187/claude-sonnet-4.6-opus-4.8-mythos-5-fable-5-openai-finetuning-dataset.5 likes873 downloads3mo agoHugging Face20lightonai /embeddings-fine-tuning-filtered-es Overview This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets. All splits except MIRACL and MLDR were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-es.3 likes842 downloads2mo agoHugging Face21thongfamilynguyen1126 /claude-sonnet-4.6-opus-4.8-mythos-5-fable-5-openai-finetuning-dataset OpenAI-Compatible Dataset Collection A collection of 29 datasets converted to OpenAI fine-tuning format ({"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}). Summary Metric Value Total Datasets 29 Total Rows ~1.5M Total Size ~1.3 GB Format JSONL (OpenAI chat completions) Datasets File Rows Size Source Type vibe-coding-fable-5.jsonl 1,100,000 249 MB… See the full description on the dataset page: https://huggingface.co/datasets/thongfamilynguyen1126/claude-sonnet-4.6-opus-4.8-mythos-5-fable-5-openai-finetuning-dataset.0 likes834 downloads2mo agoHugging Face22ravisri /BD_Finetuningimage1K<n<10K0 likes822 downloads1y agoHugging Face23lightonai /embeddings-fine-tuning-filtered-it Overview This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets. All splits except MLDR were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from embeddings-fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-it.2 likes768 downloads2mo agoHugging Face24ArchitRastogi /USCode-QAPairs-Finetuning USCode-QueryPairs Dataset This dataset contains query-answer pairs curated from the United States Code, suitable for fine-tuning any embedding model. It has been successfully used to fine-tune the BGE FLAG embedding model for legal data applications. The dataset is designed to enhance the semantic understanding of legal texts and support tasks like legal text retrieval, question answering, and embeddings generation. Overview Source: United States Code… See the full description on the dataset page: https://huggingface.co/datasets/ArchitRastogi/USCode-QAPairs-Finetuning.texttext-retrievaln<1K0 likes762 downloads2y agoHugging Face25lightonai /embeddings-fine-tuning-filtered-pt Overview This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets. All splits except MLDR were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from embeddings-fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-pt.2 likes725 downloads2mo agoHugging Face26ev-tlt /MACE_finetuning_supplementary MACE Fine-Tuning Supplementary Supplementary data and scripts for: Tompa, T. L.; Varga-Umbrich, E.; Batatia, I.; Elena, A. M.; Bernstein, N.; Csányi, G. Fine-tuning MLIP foundation models: strategies for accuracy and transferability (2026). arXiv:2606.12704. The repository contains training datasets, mace_run_train launch scripts, Slurm logs, fine-tuned model checkpoints (.model), evaluation scripts, and processed results for the paper. Benchmark systems: lithium argyrodite… See the full description on the dataset page: https://huggingface.co/datasets/ev-tlt/MACE_finetuning_supplementary.0 likes720 downloads10d agoHugging Face27lightonai /embeddings-fine-tuning-filtered-no Overview This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets. All splits were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from embeddings-fine-tuning. Translations were… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-no.2 likes695 downloads2mo agoHugging Face28omarelsherif010 /glm-ocr-bnk-finetuning GLM-OCR Fine-Tuning Pipeline Fine-tuning GLM-OCR 0.9B (CogViT encoder + GLM-0.5B decoder) for Korean financial document table recognition using LoRA via LLaMA-Factory. Performance Targets Metric Target TEDS (2-level nested) >= 90% TEDS (3-level nested) >= 85% Korean CER <= 1% Latency <= 0.5s/page Directory Structure glm_ocr_finetuning/ ├── config/ # Training/eval YAML configs │ ├── training_config.yaml #… See the full description on the dataset page: https://huggingface.co/datasets/omarelsherif010/glm-ocr-bnk-finetuning.0 likes687 downloads7mo agoHugging Face29cross-encoder /lightonai-embeddings-fine-tuning-reranked-v1 LightOn embeddings-fine-tuning, rescored with mxbai-rerank-large-v2 This dataset is a teacher-rescored version of lightonai/embeddings-fine-tuning. For every (query, candidate-document) pair in the source, we ran mixedbread-ai/mxbai-rerank-large-v2 and stored the resulting score. The point is to make the source data usable as a teacher target for distilling reranker students. It's the upstream artifact behind the rerank-scored configs of cross-encoder/ettin-reranker-v1-data… See the full description on the dataset page: https://huggingface.co/datasets/cross-encoder/lightonai-embeddings-fine-tuning-reranked-v1.texttext-ranking10M<n<100M10 likes672 downloads4mo agoHugging Face30lightonai /embeddings-fine-tuning-filtered-code Overview This dataset is composed of high quality code retrieval data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong code retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using the CoRNStack dataset. The negatives were mined following the NV-Retriever setup: the closest documents to each query are mined as negatives, and false negatives are filtered out… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-code.2 likes644 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.