datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tiny-supervised-datasetnv-embed-supervised-distill-dedup-codeThis dataset is a collection of the CoIR training datasets. We mined 2048 negatives per queries using gte-modernbert-base in order and format the data in a query, documents, scores format so that anyone can perform nv-retriever type of filtering using their own threshold (and this is also the format knowledge distillation for PyLate).
Notably, this dataset has been used to perform the fine-tuning of the state-of-the-art late interaction LateOn-Code models. The boilerplate used to fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/nv-embed-supervised-distill-dedup-code.nv-embed-supervised-distill-dedupsupervised-finetuning_quiz_student_responsesqwen3-30b-a3b-base-reasoning-sft-nemotron-math-v4-cot4k12k-500m-supervised
Qwen3-30B-A3B Reasoning SFT Prepacked Nemotron Math v4 CoT 4k-12k
This dataset is a train-ready, offline-prepacked SFT corpus for full supervised
fine-tuning of Qwen/Qwen3-30B-A3B-Base into a math reasoning model.
Source And Filtering
Source dataset: nvidia/Nemotron-SFT-Math-v4
Source revision: a94e56aeddcf6e75d28c8bd210f40fa62309288d
Source split: train
Intended subset: cot
Preferred source during selection: AoPS
Length filter: 4,000 to 12,000 supervised… See the full description on the dataset page: https://huggingface.co/datasets/ar0cket1/qwen3-30b-a3b-base-reasoning-sft-nemotron-math-v4-cot4k12k-500m-supervised.nv-embed-supervised-distillnomic-embed-supervised-dataembeddings_supervisednomic-embed-v2-supervised-dataindian-legal-supervised-fine-tuning-data
🇮🇳 LegalBrain Indic Legal Corpus
A large-scale multilingual Indian legal dataset curated to support research in:
Domain-specific LLM training
Legal question answering
Policy reasoning & case retrieval
Agentic systems for legal workflow automation
This dataset contains text drawn from publicly available legal sources across multiple Indian languages, including:
English, Hindi, Marathi, Bengali, Kannada, Tamil, Telugu, Odia, and others.
The corpus is structured and processed to be… See the full description on the dataset page: https://huggingface.co/datasets/Prarabdha/indian-legal-supervised-fine-tuning-data.supervised-multilingualmp_self_supervised
Dataset Details
Dataset Description
The materials project is a dabase of computed properties of materials.
Curated by:
License: CC BY 4.0
Dataset Sources
original data source
Citation
BibTeX:
@article{jain2013commentary,
title={Commentary: The Materials Project: A materials genome approach to accelerating materials innovation},
author={Jain, Anubhav and Ong, Shyue Ping and Hautier, Geoffroy and Chen, Wei and Richards, William Davidson and… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/mp_self_supervised.turkish_weakly_supervised_contrastive_learning_datasetnomic_embed_supervisedt5_large_supervised_proportional_1MThis data set is created by randomly sampling 1M documents from the large supervised proportional mixture from the T5 repository.
The code to produce this sampled dataset can be found here.
posttrain_tokenized_various_supervised_sup_qwen2.5_32b_instrnv-embed-supervisedsupervised_kalmvulnerable-code-snippets-for-supervised-learning
Dataset Card for vulnerable-code-snippets-for-supervised-learning
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/whackthejacker/vulnerable-code-snippets-for-supervised-learning/raw/main/pipeline.yaml"
or explore the configuration:… See the full description on the dataset page: https://huggingface.co/datasets/whackthejacker/vulnerable-code-snippets-for-supervised-learning.nomic-embed-supervised-dataSelf-Supervised_RLThis repository contains the dataset and resources related to the paper Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following.
The paper introduces a self-supervised reinforcement learning (RL) framework that improves instruction following capabilities of reasoning models by leveraging their internal signals, without requiring external supervision. This approach aims to address the trade-off between reasoning and instruction following, offering a… See the full description on the dataset page: https://huggingface.co/datasets/dd12345789/Self-Supervised_RL.LuminBench-Weakly-Supervised-CellSeg
LuminBench: Weakly Supervised Cell Segmentation
A checksum-bound training and validation package for agent-driven research with two labelled images per group. The scientific protocol and runnable code are maintained in LuminBench-Weakly-Supervised-CellSeg. The project uses the LB-Template organization for ordinary experiments and agent research.
Split
Images
Masks
Role
training_labelled
24
24
Two human-labelled examples in each of 12 groups
training_unlabelled
967
0… See the full description on the dataset page: https://huggingface.co/datasets/LuminScience/LuminBench-Weakly-Supervised-CellSeg.ViNLI-SimCSE-supervisedsupervised-finetuning_quiz_student_responsessupervised-and-blurred-rotated-shrunk-NoRevsindian-legal-supervised-fine-tuning-data
🇮🇳 LegalBrain Indic Legal Corpus
A large-scale multilingual Indian legal dataset curated to support research in:
Domain-specific LLM training
Legal question answering
Policy reasoning & case retrieval
Agentic systems for legal workflow automation
This dataset contains text drawn from publicly available legal sources across multiple Indian languages, including:
English, Hindi, Marathi, Bengali, Kannada, Tamil, Telugu, Odia, and others.
The corpus is structured and processed to be… See the full description on the dataset page: https://huggingface.co/datasets/antonhome/indian-legal-supervised-fine-tuning-data.supervisedViNLI-Healthcare-supervisedgsm8k-supervised-uncertainty-cache
GSM8K supervised uncertainty cache
Frozen feature cache used by the LM-Polygraph supervised-uncertainty seminar.
It is provided so the notebook can run without regenerating model outputs.
Contents
gsm8k_native_all_layer_compact_v1_200_100_100_c4_20_source701.joblib contains one compressed joblib payload
(835485642 bytes; SHA-256 9ed7f5beb05b6055a5a8ba04da16c086bb827b452e681f98ceee4189f8413266) with:
200 GSM8K training records and 100 disjoint GSM8K test records;… See the full description on the dataset page: https://huggingface.co/datasets/ArtemVazhentsev21/gsm8k-supervised-uncertainty-cache.Supervised-Fog-Removal-DatasetSupervised Fog Removal Dataset
Overview
This dataset contains 80,000 paired images designed for supervised image dehazing / fog removal tasks.
Each sample consists of:
a clean image (ground truth)
a synthetically fogged version of that image
The fog is generated using a physics-inspired atmospheric scattering model combined with depth estimation, allowing the fog to behave realistically with respect to scene geometry.
Unlike simple uniform haze overlays, this dataset simulates depth-aware fog… See the full description on the dataset page: https://huggingface.co/datasets/Aeye-coder/Supervised-Fog-Removal-Dataset.
