datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GIT-SCRAPED
🚀 SKT-NRS / GIT-SCRAPED
This repository is dedicated to hosting structural, curated, and diverse datasets—including GitHub roadmaps,and Roadmaps.sh Sites system architectures, technical diagrams, and mass scraped assets.
Our ultimate mission is to fuel the development of next-generation Sovereign Indian Intelligence base models with high-fidelity, production-grade text-image structures.
📂 Repository Structure
All the raw and structured crawled data is… See the full description on the dataset page: https://huggingface.co/datasets/SKT-NRS/GIT-SCRAPED.yapeichang-hotpotqa-filtered-MC_hf_qwen3_8b-NC_5-ET_0-MNT_512-NRS_3-T_1.0-S_0NR-short-32k-16-thoughts-4k-8-thoughts-Qwen3-1.7BNR-short-32k-16-thoughts-1k-8-thoughtsSubsample of JackHsieh/NR-short-32k-16-thoughts-4k-128-thoughts, with a smaller test set.
NR-short-32k-16-thoughts-4k-128-thoughtsperguntas_sft_nrs
Perguntas e Respostas SFT sobre Normas Regulamentadoras (NRs)
Dataset de pares pergunta-resposta em português (pt-BR) para supervised fine-tuning (SFT)
de modelos de linguagem no domínio de Segurança e Saúde no Trabalho (SST), com foco nas
Normas Regulamentadoras (NRs) brasileiras e em material complementar de procedimentos
internos de SESMT.
Conteúdo do dataset
Cada exemplo é um par pergunta/resposta gerado a partir de um trecho (chunk) de uma NR ou de
um… See the full description on the dataset page: https://huggingface.co/datasets/UFGCEMIGONA/perguntas_sft_nrs.nuScenes-NRS
nuScenes-NRS
nuScenes-NRS (nuScenes Nighttime Road Segmentation) is the road-segmentation
label release used by IAF-Net: Illumination-Adaptive Fusion for Low-Light
Urban Road Segmentation. This repository contains the derived labels and
the information needed to reproduce them from an authorized copy of the
official nuScenes data.
What is included
split
scenes
masks
mask resolution
distribution
training
79
3,182
1600 x 900
training/masks.zip… See the full description on the dataset page: https://huggingface.co/datasets/PeterNano/nuScenes-NRS.NR-short-32k-16-thoughts-4k-8-thoughtsSubsample of JackHsieh/NR-short-32k-16-thoughts-4k-128-thoughts, with a smaller test set.
NR-shortNRSD-MN-relabelDatasets URL:https://drive.google.com/drive/folders/13r-l_OEUt63A8K-ol6jQiaKNuGdseZ7j?usp=sharing
Datasets Paper: Chen Y, Tang Y, Hao H, et al. AMFF-YOLOX: Towards an Attention Mechanism and Multiple Feature Fusion Based on YOLOX for Industrial Defect Detection[J]. Electronics, 2023, 12(7): 1662.
Dataset Original Repository: MCnet
Dataset Original Paper: Zhang D, Song K, Xu J, et al. MCnet: Multiple context information segmentation network of no-service rail surface defects[J]. IEEE… See the full description on the dataset page: https://huggingface.co/datasets/chairc/NRSD-MN-relabel.NRS-DOCSNR-short-128-128NR-short-1k-1-thought-1k-1-thoughtSubsample of JackHsieh/NR-short-32k-16-thoughts-4k-128-thoughts, with a much smaller test set.
NR-short-32k-1kJackHsieh/NR-short-32k-16-thoughts-1k-8-thoughts, stripped of thoughts and deduplicated.
NR-short-1k-1kJackHsieh/NR-short-1k-1-thought-1k-1-thought, stripped of thoughts.
NR-short-32k-4kST-CORE-TOKENS
🇮🇳 SOVEREIGN INDIAN INTELLIGENCE 🇮🇳
🔥 SKT AI LABS 🔥
SKT AI LABS
The Sovereign AI for India
🇮🇳 MADE IN BHARAT
🧬 ST-CORE-TOKENS
LOGIC DISTILLED
588 GB+
🔱 SKT-Ai-Labs/ST-CORE-TOKENS
ST-CORE-TOKENS ek ultra-refined, high-density tokenized dataset hai jo SKT AI LABS dwara develop kiya gaya hai. Yeh dataset Indian LLMs ki Cognitive Logic… See the full description on the dataset page: https://huggingface.co/datasets/SKT-NRS/ST-CORE-TOKENS.ST-CODEX
ST-CODEX
AUDIO-SRX-OTHER-LANGUAGE
SKT AI LABS
SKT AI LABS
The Sovereign AI for India
The Sovereign LLM Development For India (Project Surya)
SKT-OMNI-CORPUS-2T
SKT AI LABS INDIA
PROJECT SURYA
INDIA'S SOVEREIGN AI FOUNDATION
🔱 SKT-OMNI-CORPUS-2T | 5TB • Building Towards 8TB
The first truly sovereign AI foundation model built on native Hindi & Hinglish data.
Processing the thoughts of 1.4 billion people — from Sidhi to Siliguri.
🚀 Welcome to Project Surya (सूर्य)
Open Sourced Now! Project Surya is not just another LLM — it is India's first… See the full description on the dataset page: https://huggingface.co/datasets/SKT-NRS/SKT-OMNI-CORPUS-2T.ST-TRAIN
🚂 SOVEREIGN TRAINING CORPUS 🚂
🔥 ST-TRAIN 🔥
FOUNDATION CORPUS
ST-V2
MASSIVE SCALE
🔱 SKT-Ai-Labs/ST-TRAIN
ST-TRAIN ek massive, high-quality training corpus hai jo SKT AI LABS dwara Sovereign Indian LLMs ki pre-training ke liye curate kiya gaya hai. Ye dataset foundation models ko Bharat ki linguistic aur cultural nuances sikhane ka primary source hai.
🛰️ Key Features
Massive Diversity: General knowledge se lekar specific technical… See the full description on the dataset page: https://huggingface.co/datasets/SKT-NRS/ST-TRAIN.WIMGnrsSKT-TOKENS
🇮🇳 SOVEREIGN INDIAN INTELLIGENCE 🇮🇳
🔥 SKT AI LABS 🔥
SKT AI LABS
The Sovereign AI for India
🇮🇳 MADE IN BHARAT
🧬 SKT-TOKENS
LOGIC CORE
500 GB+
Knowlege Booster
🔱 SKT-Ai-Labs/SKT-TOKENS
SKT-TOKENS ek high-density, logic-distilled dataset hai jo SKT AI LABS dwara "Project Surya" aur anya Sovereign Indian LLMs ke liye curate kiya gaya hai. Isme… See the full description on the dataset page: https://huggingface.co/datasets/SKT-NRS/SKT-TOKENS.NRS
🧠 NEURAL REASONING SYSTEM 🧠
⚡ NRS CORE ⚡
COGNITIVE LOGIC
NRS-V1
REASONING ENGINE
🔱 SKT-Ai-Labs/NRS
NRS (Neural Reasoning System) ek specialized dataset hai jise SKT AI LABS ne develop kiya hai. Iska main focus LLMs ke andar "Chain-of-Thought" (CoT) aur complex step-by-step reasoning capabilities ko build karna hai.
🛰️ Key Features
Deep Reasoning: Multi-step math, coding, aur logical puzzles ka collection.
Thought Traces: Model ko… See the full description on the dataset page: https://huggingface.co/datasets/SKT-NRS/NRS.yapeichang-hotpotqa-filtered-MC_hf_qwen3_8b-NC_5-ET_0-MNT_512-NRS_3-T_1.0-S_42ST-TOKENS
🇮🇳 SOVEREIGN INDIAN INTELLIGENCE 🇮🇳
🔥 SKT AI LABS 🔥
SKT AI LABS
The Sovereign AI for India
🇮🇳 MADE IN BHARAT
🧬 ST-CORE-TOKENS
PURE LOGIC DISTILLED
600 GB+
Booster
---
🔱 SKT-Ai-Labs/ST-TOKENS
ST-TOKENS ek ultra-massive, high-density logic-distilled dataset hai jise SKT AI LABS ne develop kiya hai. Yeh Bharat ke "Sovereign AI"… See the full description on the dataset page: https://huggingface.co/datasets/SKT-NRS/ST-TOKENS.Anadi-nrscnRSM3BeUNRSR_energetika
