CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lightonai /nv-embed-supervised-distill-dedup-codeThis dataset is a collection of the CoIR training datasets. We mined 2048 negatives per queries using gte-modernbert-base in order and format the data in a query, documents, scores format so that anyone can perform nv-retriever type of filtering using their own threshold (and this is also the format knowledge distillation for PyLate). Notably, this dataset has been used to perform the fine-tuning of the state-of-the-art late interaction LateOn-Code models. The boilerplate used to fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/nv-embed-supervised-distill-dedup-code.text1M<n<10M7 likes1.7k downloads3mo agoHugging Face02lightonai /nv-embed-supervised-distill-deduptext10M<n<100M0 likes1.6k downloads5mo agoHugging Face03huggingface-course /supervised-finetuning_quiz_student_responsestextn<1K4 likes1.2k downloads2h agoHugging Face04lightonai /nv-embed-supervised-distilltext10M<n<100M1 likes996 downloads11mo agoHugging Face05nomic-ai /nomic-embed-supervised-datatext1M<n<10M6 likes675 downloads2y agoHugging Face06Prarabdha /indian-legal-supervised-fine-tuning-data 🇮🇳 LegalBrain Indic Legal Corpus A large-scale multilingual Indian legal dataset curated to support research in: Domain-specific LLM training Legal question answering Policy reasoning & case retrieval Agentic systems for legal workflow automation This dataset contains text drawn from publicly available legal sources across multiple Indian languages, including: English, Hindi, Marathi, Bengali, Kannada, Tamil, Telugu, Odia, and others. The corpus is structured and processed to be… See the full description on the dataset page: https://huggingface.co/datasets/Prarabdha/indian-legal-supervised-fine-tuning-data.text1M<n<10M8 likes517 downloads11mo agoHugging Face07lightonai /embeddings_supervisedtext1M<n<10M13 likes480 downloads11mo agoHugging Face08nomic-ai /nomic-embed-v2-supervised-datatext1M<n<10M2 likes458 downloads1y agoHugging Face09gowitheflow /supervised-multilingualtext10M<n<100M1 likes296 downloads2y agoHugging Face10jablonkagroup /mp_self_supervised Dataset Details Dataset Description The materials project is a dabase of computed properties of materials. Curated by: License: CC BY 4.0 Dataset Sources original data source Citation BibTeX: @article{jain2013commentary, title={Commentary: The Materials Project: A materials genome approach to accelerating materials innovation}, author={Jain, Anubhav and Ong, Shyue Ping and Hautier, Geoffroy and Chen, Wei and Richards, William Davidson and… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/mp_self_supervised.tabular1M<n<10M0 likes253 downloads1y agoHugging Face11selmanbaysan /turkish_weakly_supervised_contrastive_learning_datasettext10M<n<100M0 likes198 downloads1y agoHugging Face12jxm /nomic_embed_supervisedtext1M<n<10M2 likes163 downloads2y agoHugging Face13ZixuanKe /posttrain_tokenized_various_supervised_sup_qwen2.5_32b_instrtext1M<n<10M0 likes124 downloads2y agoHugging Face14lightonai /nv-embed-supervisedtext100K<n<1M2 likes118 downloads11mo agoHugging Face15lightonai /supervised_kalmtext1M<n<10M5 likes101 downloads11mo agoHugging Face16whackthejacker /vulnerable-code-snippets-for-supervised-learning Dataset Card for vulnerable-code-snippets-for-supervised-learning This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/whackthejacker/vulnerable-code-snippets-for-supervised-learning/raw/main/pipeline.yaml" or explore the configuration:… See the full description on the dataset page: https://huggingface.co/datasets/whackthejacker/vulnerable-code-snippets-for-supervised-learning.texttext-classificationn<1K0 likes95 downloads2y agoHugging Face17raul3820 /nomic-embed-supervised-datatext1M<n<10M0 likes88 downloads6mo agoHugging Face18dd12345789 /Self-Supervised_RLThis repository contains the dataset and resources related to the paper Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following. The paper introduces a self-supervised reinforcement learning (RL) framework that improves instruction following capabilities of reasoning models by leveraging their internal signals, without requiring external supervision. This approach aims to address the trade-off between reasoning and instruction following, offering a… See the full description on the dataset page: https://huggingface.co/datasets/dd12345789/Self-Supervised_RL.texttext-generation10K<n<100K0 likes83 downloads5mo agoHugging Face19jealk /supervised-datext100K<n<1M1 likes59 downloads2y agoHugging Face20reasoning-course /supervised-finetuning_quiz_student_responsestextn<1K2 likes51 downloads2y agoHugging Face21simondarius /supervisedtimeseries10K<n<100K0 likes40 downloads2y agoHugging Face22kahua-ml /supervised-and-blurred-rotated-shrunk-NoRevsimagen<1K1 likes38 downloads1y agoHugging Face23antonhome /indian-legal-supervised-fine-tuning-data 🇮🇳 LegalBrain Indic Legal Corpus A large-scale multilingual Indian legal dataset curated to support research in: Domain-specific LLM training Legal question answering Policy reasoning & case retrieval Agentic systems for legal workflow automation This dataset contains text drawn from publicly available legal sources across multiple Indian languages, including: English, Hindi, Marathi, Bengali, Kannada, Tamil, Telugu, Odia, and others. The corpus is structured and processed to be… See the full description on the dataset page: https://huggingface.co/datasets/antonhome/indian-legal-supervised-fine-tuning-data.text1M<n<10M0 likes38 downloads9mo agoHugging Face24Aeye-coder /Supervised-Fog-Removal-DatasetSupervised Fog Removal Dataset Overview This dataset contains 80,000 paired images designed for supervised image dehazing / fog removal tasks. Each sample consists of: a clean image (ground truth) a synthetically fogged version of that image The fog is generated using a physics-inspired atmospheric scattering model combined with depth estimation, allowing the fog to behave realistically with respect to scene geometry. Unlike simple uniform haze overlays, this dataset simulates depth-aware fog… See the full description on the dataset page: https://huggingface.co/datasets/Aeye-coder/Supervised-Fog-Removal-Dataset.image10K<n<100K1 likes30 downloads4mo agoHugging Face25selmanbaysan /turkish_weakly_supervised_contrastive_learning_dataset_filteredtabular1K<n<10K0 likes24 downloads1y agoHugging Face26sibasmarakp /Qwen2.5-Math-7B-Instruct-Qwen2.5-14B-Instruct-SupervisedPRM-T80-adapters-best_of_n-completionstabular1K<n<10K0 likes22 downloads5mo agoHugging Face27BookingCare /ViNLI-SimCSE-supervisedtext100K<n<1M0 likes17 downloads2y agoHugging Face28simondarius /MiniWoB_supervised_pretrainingtimeseries1K<n<10K0 likes16 downloads2y agoHugging Face29Triptigarg2711 /supervised_finetuning2text1K<n<10K0 likes16 downloads2y agoHugging Face30sibasmarakp /Qwen2.5-Math-7B-Instruct-SupervisedPRM-T80-adapters-best_of_n-completionstabular1K<n<10K0 likes14 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.