CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lightonai /nv-embed-supervised-distill-deduptext10M<n<100M0 likes2k downloads5mo agoHugging Face02lightonai /nv-embed-supervised-distill-dedup-codeThis dataset is a collection of the CoIR training datasets. We mined 2048 negatives per queries using gte-modernbert-base in order and format the data in a query, documents, scores format so that anyone can perform nv-retriever type of filtering using their own threshold (and this is also the format knowledge distillation for PyLate). Notably, this dataset has been used to perform the fine-tuning of the state-of-the-art late interaction LateOn-Code models. The boilerplate used to fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/nv-embed-supervised-distill-dedup-code.text1M<n<10M7 likes1.8k downloads3mo agoHugging Face03lightonai /nv-embed-supervised-distilltext10M<n<100M1 likes1.3k downloads11mo agoHugging Face04huggingface-course /supervised-finetuning_quiz_student_responsestextn<1K4 likes1.2k downloads9h agoHugging Face05nomic-ai /nomic-embed-supervised-datatext1M<n<10M6 likes576 downloads2y agoHugging Face06Prarabdha /indian-legal-supervised-fine-tuning-data 🇮🇳 LegalBrain Indic Legal Corpus A large-scale multilingual Indian legal dataset curated to support research in: Domain-specific LLM training Legal question answering Policy reasoning & case retrieval Agentic systems for legal workflow automation This dataset contains text drawn from publicly available legal sources across multiple Indian languages, including: English, Hindi, Marathi, Bengali, Kannada, Tamil, Telugu, Odia, and others. The corpus is structured and processed to be… See the full description on the dataset page: https://huggingface.co/datasets/Prarabdha/indian-legal-supervised-fine-tuning-data.text1M<n<10M8 likes524 downloads11mo agoHugging Face07lightonai /embeddings_supervisedtext1M<n<10M13 likes521 downloads11mo agoHugging Face08nomic-ai /nomic-embed-v2-supervised-datatext1M<n<10M2 likes418 downloads1y agoHugging Face09gowitheflow /supervised-multilingualtext10M<n<100M1 likes314 downloads2y agoHugging Face10jablonkagroup /mp_self_supervised Dataset Details Dataset Description The materials project is a dabase of computed properties of materials. Curated by: License: CC BY 4.0 Dataset Sources original data source Citation BibTeX: @article{jain2013commentary, title={Commentary: The Materials Project: A materials genome approach to accelerating materials innovation}, author={Jain, Anubhav and Ong, Shyue Ping and Hautier, Geoffroy and Chen, Wei and Richards, William Davidson and… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/mp_self_supervised.tabular1M<n<10M0 likes305 downloads1y agoHugging Face11selmanbaysan /turkish_weakly_supervised_contrastive_learning_datasettext10M<n<100M0 likes175 downloads1y agoHugging Face12jxm /nomic_embed_supervisedtext1M<n<10M2 likes160 downloads2y agoHugging Face13ZixuanKe /posttrain_tokenized_various_supervised_sup_qwen2.5_32b_instrtext1M<n<10M0 likes155 downloads2y agoHugging Face14lightonai /nv-embed-supervisedtext100K<n<1M2 likes127 downloads11mo agoHugging Face15lightonai /supervised_kalmtext1M<n<10M5 likes110 downloads11mo agoHugging Face16whackthejacker /vulnerable-code-snippets-for-supervised-learning Dataset Card for vulnerable-code-snippets-for-supervised-learning This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/whackthejacker/vulnerable-code-snippets-for-supervised-learning/raw/main/pipeline.yaml" or explore the configuration:… See the full description on the dataset page: https://huggingface.co/datasets/whackthejacker/vulnerable-code-snippets-for-supervised-learning.texttext-classificationn<1K0 likes91 downloads2y agoHugging Face17raul3820 /nomic-embed-supervised-datatext1M<n<10M0 likes87 downloads6mo agoHugging Face18dd12345789 /Self-Supervised_RLThis repository contains the dataset and resources related to the paper Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following. The paper introduces a self-supervised reinforcement learning (RL) framework that improves instruction following capabilities of reasoning models by leveraging their internal signals, without requiring external supervision. This approach aims to address the trade-off between reasoning and instruction following, offering a… See the full description on the dataset page: https://huggingface.co/datasets/dd12345789/Self-Supervised_RL.texttext-generation10K<n<100K0 likes86 downloads5mo agoHugging Face19jealk /supervised-datext100K<n<1M1 likes60 downloads2y agoHugging Face20reasoning-course /supervised-finetuning_quiz_student_responsestextn<1K2 likes50 downloads2y agoHugging Face21kahua-ml /supervised-and-blurred-rotated-shrunk-NoRevsimagen<1K1 likes40 downloads2y agoHugging Face22antonhome /indian-legal-supervised-fine-tuning-data 🇮🇳 LegalBrain Indic Legal Corpus A large-scale multilingual Indian legal dataset curated to support research in: Domain-specific LLM training Legal question answering Policy reasoning & case retrieval Agentic systems for legal workflow automation This dataset contains text drawn from publicly available legal sources across multiple Indian languages, including: English, Hindi, Marathi, Bengali, Kannada, Tamil, Telugu, Odia, and others. The corpus is structured and processed to be… See the full description on the dataset page: https://huggingface.co/datasets/antonhome/indian-legal-supervised-fine-tuning-data.text1M<n<10M0 likes37 downloads9mo agoHugging Face23selmanbaysan /turkish_weakly_supervised_contrastive_learning_dataset_filteredtabular1K<n<10K0 likes24 downloads1y agoHugging Face24sibasmarakp /Qwen2.5-Math-7B-Instruct-Qwen2.5-14B-Instruct-SupervisedPRM-T80-adapters-best_of_n-completionstabular1K<n<10K0 likes20 downloads5mo agoHugging Face25BookingCare /ViNLI-SimCSE-supervisedtext100K<n<1M0 likes17 downloads2y agoHugging Face26Triptigarg2711 /supervised_finetuning2text1K<n<10K0 likes15 downloads2y agoHugging Face27sibasmarakp /Qwen2.5-Math-7B-Instruct-SupervisedPRM-T80-adapters-best_of_n-completionstabular1K<n<10K0 likes13 downloads6mo agoHugging Face28li-ping /1127_supervised_ft_embedding_v1textn<1K0 likes12 downloads3y agoHugging Face29illuin-conteb /nomic_embed_supervised_clusteredtext1M<n<10M0 likes10 downloads2y agoHugging Face30harshit03 /supervisedDatasettextn<1K0 likes8 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.