CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01llamafactory /tiny-supervised-datasettexttext-generationn<1K4 likes43k downloads2y agoHugging Face02lightonai /nv-embed-supervised-distill-dedup-codeThis dataset is a collection of the CoIR training datasets. We mined 2048 negatives per queries using gte-modernbert-base in order and format the data in a query, documents, scores format so that anyone can perform nv-retriever type of filtering using their own threshold (and this is also the format knowledge distillation for PyLate). Notably, this dataset has been used to perform the fine-tuning of the state-of-the-art late interaction LateOn-Code models. The boilerplate used to fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/nv-embed-supervised-distill-dedup-code.text1M<n<10M7 likes1.7k downloads2mo agoHugging Face03lightonai /nv-embed-supervised-distill-deduptext10M<n<100M0 likes1.6k downloads5mo agoHugging Face04huggingface-course /supervised-finetuning_quiz_student_responses4 likes1.2k downloads1h agoHugging Face05ar0cket1 /qwen3-30b-a3b-base-reasoning-sft-nemotron-math-v4-cot4k12k-500m-supervised Qwen3-30B-A3B Reasoning SFT Prepacked Nemotron Math v4 CoT 4k-12k This dataset is a train-ready, offline-prepacked SFT corpus for full supervised fine-tuning of Qwen/Qwen3-30B-A3B-Base into a math reasoning model. Source And Filtering Source dataset: nvidia/Nemotron-SFT-Math-v4 Source revision: a94e56aeddcf6e75d28c8bd210f40fa62309288d Source split: train Intended subset: cot Preferred source during selection: AoPS Length filter: 4,000 to 12,000 supervised… See the full description on the dataset page: https://huggingface.co/datasets/ar0cket1/qwen3-30b-a3b-base-reasoning-sft-nemotron-math-v4-cot4k12k-500m-supervised.1 likes1.2k downloads2mo agoHugging Face06lightonai /nv-embed-supervised-distilltext10M<n<100M1 likes997 downloads11mo agoHugging Face07nomic-ai /nomic-embed-supervised-datatext1M<n<10M6 likes704 downloads2y agoHugging Face08lightonai /embeddings_supervisedtext1M<n<10M13 likes478 downloads11mo agoHugging Face09nomic-ai /nomic-embed-v2-supervised-datatext1M<n<10M2 likes471 downloads1y agoHugging Face10Prarabdha /indian-legal-supervised-fine-tuning-data 🇮🇳 LegalBrain Indic Legal Corpus A large-scale multilingual Indian legal dataset curated to support research in: Domain-specific LLM training Legal question answering Policy reasoning & case retrieval Agentic systems for legal workflow automation This dataset contains text drawn from publicly available legal sources across multiple Indian languages, including: English, Hindi, Marathi, Bengali, Kannada, Tamil, Telugu, Odia, and others. The corpus is structured and processed to be… See the full description on the dataset page: https://huggingface.co/datasets/Prarabdha/indian-legal-supervised-fine-tuning-data.text1M<n<10M8 likes453 downloads11mo agoHugging Face11gowitheflow /supervised-multilingualtext10M<n<100M1 likes297 downloads2y agoHugging Face12jablonkagroup /mp_self_supervised Dataset Details Dataset Description The materials project is a dabase of computed properties of materials. Curated by: License: CC BY 4.0 Dataset Sources original data source Citation BibTeX: @article{jain2013commentary, title={Commentary: The Materials Project: A materials genome approach to accelerating materials innovation}, author={Jain, Anubhav and Ong, Shyue Ping and Hautier, Geoffroy and Chen, Wei and Richards, William Davidson and… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/mp_self_supervised.tabular1M<n<10M0 likes243 downloads1y agoHugging Face13selmanbaysan /turkish_weakly_supervised_contrastive_learning_datasettext10M<n<100M0 likes193 downloads1y agoHugging Face14jxm /nomic_embed_supervisedtext1M<n<10M2 likes162 downloads2y agoHugging Face15jchenyu /t5_large_supervised_proportional_1MThis data set is created by randomly sampling 1M documents from the large supervised proportional mixture from the T5 repository. The code to produce this sampled dataset can be found here. tabular1M<n<10M0 likes140 downloads4y agoHugging Face16ZixuanKe /posttrain_tokenized_various_supervised_sup_qwen2.5_32b_instrtext1M<n<10M0 likes124 downloads2y agoHugging Face17lightonai /nv-embed-supervisedtext100K<n<1M2 likes118 downloads11mo agoHugging Face18lightonai /supervised_kalmtext1M<n<10M5 likes103 downloads11mo agoHugging Face19whackthejacker /vulnerable-code-snippets-for-supervised-learning Dataset Card for vulnerable-code-snippets-for-supervised-learning This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/whackthejacker/vulnerable-code-snippets-for-supervised-learning/raw/main/pipeline.yaml" or explore the configuration:… See the full description on the dataset page: https://huggingface.co/datasets/whackthejacker/vulnerable-code-snippets-for-supervised-learning.texttext-classificationn<1K0 likes95 downloads2y agoHugging Face20raul3820 /nomic-embed-supervised-datatext1M<n<10M0 likes88 downloads6mo agoHugging Face21dd12345789 /Self-Supervised_RLThis repository contains the dataset and resources related to the paper Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following. The paper introduces a self-supervised reinforcement learning (RL) framework that improves instruction following capabilities of reasoning models by leveraging their internal signals, without requiring external supervision. This approach aims to address the trade-off between reasoning and instruction following, offering a… See the full description on the dataset page: https://huggingface.co/datasets/dd12345789/Self-Supervised_RL.texttext-generation10K<n<100K0 likes79 downloads5mo agoHugging Face22LuminScience /LuminBench-Weakly-Supervised-CellSeg LuminBench: Weakly Supervised Cell Segmentation A checksum-bound training and validation package for agent-driven research with two labelled images per group. The scientific protocol and runnable code are maintained in LuminBench-Weakly-Supervised-CellSeg. The project uses the LB-Template organization for ordinary experiments and agent research. Split Images Masks Role training_labelled 24 24 Two human-labelled examples in each of 12 groups training_unlabelled 967 0… See the full description on the dataset page: https://huggingface.co/datasets/LuminScience/LuminBench-Weakly-Supervised-CellSeg.imageimage-segmentation1K<n<10K0 likes62 downloads5d agoHugging Face23anti-ai /ViNLI-SimCSE-supervisedtextsentence-similarity100K<n<1M1 likes51 downloads3y agoHugging Face24reasoning-course /supervised-finetuning_quiz_student_responsestextn<1K2 likes50 downloads2y agoHugging Face25kahua-ml /supervised-and-blurred-rotated-shrunk-NoRevsimagen<1K1 likes38 downloads1y agoHugging Face26antonhome /indian-legal-supervised-fine-tuning-data 🇮🇳 LegalBrain Indic Legal Corpus A large-scale multilingual Indian legal dataset curated to support research in: Domain-specific LLM training Legal question answering Policy reasoning & case retrieval Agentic systems for legal workflow automation This dataset contains text drawn from publicly available legal sources across multiple Indian languages, including: English, Hindi, Marathi, Bengali, Kannada, Tamil, Telugu, Odia, and others. The corpus is structured and processed to be… See the full description on the dataset page: https://huggingface.co/datasets/antonhome/indian-legal-supervised-fine-tuning-data.text1M<n<10M0 likes38 downloads9mo agoHugging Face27simondarius /supervisedtimeseries10K<n<100K0 likes36 downloads2y agoHugging Face28anti-ai /ViNLI-Healthcare-supervisedgatedtexttext-classification1M<n<10M0 likes35 downloads1y agoHugging Face29ArtemVazhentsev21 /gsm8k-supervised-uncertainty-cache GSM8K supervised uncertainty cache Frozen feature cache used by the LM-Polygraph supervised-uncertainty seminar. It is provided so the notebook can run without regenerating model outputs. Contents gsm8k_native_all_layer_compact_v1_200_100_100_c4_20_source701.joblib contains one compressed joblib payload (835485642 bytes; SHA-256 9ed7f5beb05b6055a5a8ba04da16c086bb827b452e681f98ceee4189f8413266) with: 200 GSM8K training records and 100 disjoint GSM8K test records;… See the full description on the dataset page: https://huggingface.co/datasets/ArtemVazhentsev21/gsm8k-supervised-uncertainty-cache.text-generationn<1K0 likes33 downloads1mo agoHugging Face30Aeye-coder /Supervised-Fog-Removal-DatasetSupervised Fog Removal Dataset Overview This dataset contains 80,000 paired images designed for supervised image dehazing / fog removal tasks. Each sample consists of: a clean image (ground truth) a synthetically fogged version of that image The fog is generated using a physics-inspired atmospheric scattering model combined with depth estimation, allowing the fog to behave realistically with respect to scene geometry. Unlike simple uniform haze overlays, this dataset simulates depth-aware fog… See the full description on the dataset page: https://huggingface.co/datasets/Aeye-coder/Supervised-Fog-Removal-Dataset.image10K<n<100K1 likes30 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.