datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nv-embed-supervised-distill-dedup-codeThis dataset is a collection of the CoIR training datasets. We mined 2048 negatives per queries using gte-modernbert-base in order and format the data in a query, documents, scores format so that anyone can perform nv-retriever type of filtering using their own threshold (and this is also the format knowledge distillation for PyLate).
Notably, this dataset has been used to perform the fine-tuning of the state-of-the-art late interaction LateOn-Code models. The boilerplate used to fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/nv-embed-supervised-distill-dedup-code.nv-embed-supervised-distill-dedupsupervised-finetuning_quiz_student_responsesnv-embed-supervised-distillnomic-embed-supervised-dataindian-legal-supervised-fine-tuning-data
🇮🇳 LegalBrain Indic Legal Corpus
A large-scale multilingual Indian legal dataset curated to support research in:
Domain-specific LLM training
Legal question answering
Policy reasoning & case retrieval
Agentic systems for legal workflow automation
This dataset contains text drawn from publicly available legal sources across multiple Indian languages, including:
English, Hindi, Marathi, Bengali, Kannada, Tamil, Telugu, Odia, and others.
The corpus is structured and processed to be… See the full description on the dataset page: https://huggingface.co/datasets/Prarabdha/indian-legal-supervised-fine-tuning-data.embeddings_supervisednomic-embed-v2-supervised-datasupervised-multilingualmp_self_supervised
Dataset Details
Dataset Description
The materials project is a dabase of computed properties of materials.
Curated by:
License: CC BY 4.0
Dataset Sources
original data source
Citation
BibTeX:
@article{jain2013commentary,
title={Commentary: The Materials Project: A materials genome approach to accelerating materials innovation},
author={Jain, Anubhav and Ong, Shyue Ping and Hautier, Geoffroy and Chen, Wei and Richards, William Davidson and… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/mp_self_supervised.turkish_weakly_supervised_contrastive_learning_datasetnomic_embed_supervisedposttrain_tokenized_various_supervised_sup_qwen2.5_32b_instrnv-embed-supervisedsupervised_kalmvulnerable-code-snippets-for-supervised-learning
Dataset Card for vulnerable-code-snippets-for-supervised-learning
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/whackthejacker/vulnerable-code-snippets-for-supervised-learning/raw/main/pipeline.yaml"
or explore the configuration:… See the full description on the dataset page: https://huggingface.co/datasets/whackthejacker/vulnerable-code-snippets-for-supervised-learning.nomic-embed-supervised-dataSelf-Supervised_RLThis repository contains the dataset and resources related to the paper Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following.
The paper introduces a self-supervised reinforcement learning (RL) framework that improves instruction following capabilities of reasoning models by leveraging their internal signals, without requiring external supervision. This approach aims to address the trade-off between reasoning and instruction following, offering a… See the full description on the dataset page: https://huggingface.co/datasets/dd12345789/Self-Supervised_RL.supervised-dasupervised-finetuning_quiz_student_responsessupervisedsupervised-and-blurred-rotated-shrunk-NoRevsindian-legal-supervised-fine-tuning-data
🇮🇳 LegalBrain Indic Legal Corpus
A large-scale multilingual Indian legal dataset curated to support research in:
Domain-specific LLM training
Legal question answering
Policy reasoning & case retrieval
Agentic systems for legal workflow automation
This dataset contains text drawn from publicly available legal sources across multiple Indian languages, including:
English, Hindi, Marathi, Bengali, Kannada, Tamil, Telugu, Odia, and others.
The corpus is structured and processed to be… See the full description on the dataset page: https://huggingface.co/datasets/antonhome/indian-legal-supervised-fine-tuning-data.Supervised-Fog-Removal-DatasetSupervised Fog Removal Dataset
Overview
This dataset contains 80,000 paired images designed for supervised image dehazing / fog removal tasks.
Each sample consists of:
a clean image (ground truth)
a synthetically fogged version of that image
The fog is generated using a physics-inspired atmospheric scattering model combined with depth estimation, allowing the fog to behave realistically with respect to scene geometry.
Unlike simple uniform haze overlays, this dataset simulates depth-aware fog… See the full description on the dataset page: https://huggingface.co/datasets/Aeye-coder/Supervised-Fog-Removal-Dataset.turkish_weakly_supervised_contrastive_learning_dataset_filteredQwen2.5-Math-7B-Instruct-Qwen2.5-14B-Instruct-SupervisedPRM-T80-adapters-best_of_n-completionsViNLI-SimCSE-supervisedMiniWoB_supervised_pretrainingsupervised_finetuning2Qwen2.5-Math-7B-Instruct-SupervisedPRM-T80-adapters-best_of_n-completions
