datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indian-government-schemes-2025
Indian Government Schemes Dataset 2026
Dataset Description
The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields.
Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India
This dataset powers SchemeFit — India's government scheme finder for citizens and businesses.
What Makes This Different
Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/smartduketech/indian-government-schemes-2025.hse-qa-corpus
Canonical landing page: https://www.smartqhse.com/datasets/hse-qa-corpus
SmartQHSE HSE Q&A Corpus
24 long-form HSE / occupational-safety question-answer pairs across 15 categories — incident rates, ISO 45001, permits, risk assessment, OSHA (US), HSE (UK), GCC regulations, PPE, heat stress, exposure, ergonomics, incident investigation, training, HSE software. Each answer is multi-paragraph with cited sources, formulas, and OSHA/regulatory references. Suitable for instruction… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/hse-qa-corpus.hse-instruction-tuning
SmartQHSE HSE Instruction-Tuning Corpus
Alpaca-style instruction-tuning variant of the SmartQHSE HSE Q&A Corpus.
Each row has the canonical fine-tuning schema:
{
"instruction": "What is OSHA Process Safety Management 1910.119?",
"input": "",
"output": "<authoritative long-form answer with citations>",
"category": "us-osha",
"source_url": "https://www.smartqhse.com/answers/<slug>"
}
Suitable for LoRA / SFT training of HSE-domain LLMs and RAG systems.
Citation (preferred… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/hse-instruction-tuning.major-process-safety-incidents-2026
Canonical landing page: https://www.smartqhse.com/datasets/major-process-safety-incidents-2026
Major Process Safety Incidents Reference Database 2026
15 catastrophic process-safety accidents from 1984-2022 — Bhopal, Piper Alpha, BP Texas City, Macondo, Buncefield, Imperial Sugar, West Fertilizer, Philadelphia Energy Solutions, BP Husky Toledo, AZF Toulouse, TVA Kingston, ExxonMobil Beaumont, Chevron Richmond, Esso Longford, Phillips 66 Pasadena. Each entry: cause, fatalities… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/major-process-safety-incidents-2026.omnimcp_smartenergy_iot_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_smartenergy_iot_teaser.reddit_dataset_239
Bittensor Subnet 13 Reddit Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more… See the full description on the dataset page: https://huggingface.co/datasets/smartnuel87/reddit_dataset_239.smart-contracts-instructions
Smart Contracts Instructions
A dataset containing 6,003 GPT-generated human instruction and Solidity source code data pairs.
GPT models used to make this data are GPT-3.5 turbo, GPT-3.5 turbo 16k context, and GPT-4. Solidity source codes are used from mwritescode's Slither Audited Smart Contracts (https://huggingface.co/datasets/mwritescode/slither-audited-smart-contracts).
Distributions of the GPT models used to make this dataset:
GPT-3.5 Turbo: 5,276
GPT-3.5 Turbo 16k Context:… See the full description on the dataset page: https://huggingface.co/datasets/AlfredPros/smart-contracts-instructions.lexic-ai-tutorial-datasetethereum-smart-contract-security-qa
Ethereum Smart Contract Security Dataset
A high-quality question–answer dataset of 100 records focused exclusively on Ethereum
smart contract security. It is built to train and evaluate AI systems that explain, identify,
classify, and mitigate the most common Ethereum smart contract vulnerabilities — LLM
fine-tuning, retrieval-augmented generation (RAG), AI security assistants, and secure
Solidity education.
Every record explains one vulnerability, attack pattern, secure coding… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/ethereum-smart-contract-security-qa.x_dataset_239
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/smartnuel87/x_dataset_239.SMART
SMART: Evaluating LLMs’ Mathematical Reasoning via a Human Cognitive Process-Inspired Benchmark
SMART is a fine-grained benchmark for evaluating large language models (LLMs) on mathematical reasoning from a human cognitive process perspective. Instead of evaluating only the final answer, SMART decomposes mathematical problem solving into four cognitive dimensions inspired by Pólya’s problem-solving theory:
Semantic Understanding
Mathematical Reasoning
Arithmetic Computation… See the full description on the dataset page: https://huggingface.co/datasets/ewdfd/SMART.squad_sr
Dataset Card for Serbian SQuAD
Dataset Summary
This dataset is an automatic Serbian translation of the Stanford Question Answering Dataset (SQuAD) 1.1. The original SQuAD, developed by Stanford, is a reading comprehension dataset consisting of questions posed by crowdworkers on a set of Wikipedia articles. The answer to every question is a segment of text (span) from the corresponding reading passage, or the question might be unanswerable. It's the largest Serbian QA… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/squad_sr.ms_marco_sr
Dataset Card for Serbian MS MARCO (Subset)
Dataset Summary
This dataset is a Serbian translation of the first 8,000 examples from Microsoft's MS MARCO (Machine Reading Comprehension) dataset. It contains pairs of questions and human-generated answers, automatically translated from English to Serbian. The dataset is designed for evaluating embedding models on Question Answering (QA) and Information Retrieval (IR) tasks in the Serbian language.
The original MS MARCO dataset… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/ms_marco_sr.natural_quesions_sr
Dataset Card for Serbian Natural Questions (Subset)
Dataset Summary
This dataset is a Serbian translation of the first 8,000 examples from Google's Natural Questions (NQ) dataset. It contains real user questions and corresponding Wikipedia articles, automatically translated from English to Serbian. The dataset is designed for evaluating embedding models on Question Answering (QA) and Information Retrieval (IR) tasks in the Serbian language, offering a more realistic and… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/natural_quesions_sr.smartroute-rag-synthetic-routing-benchmark-5000
🧭 SmartRoute-RAG Synthetic Routing Benchmark 5000
A publication-scale benchmark for evaluating when to retrieve — not just what to answer.
5,000 stratified questions · 10 benchmark-style subsets · 13 question types · binary routing labelsBuilt for the SmartRoute-RAG research line: false-skip-aware, safety-constrained adaptive retrieval.
🎯 Why this dataset exists
Most RAG benchmarks measure answer quality after retrieval. They rarely tell you whether the… See the full description on the dataset page: https://huggingface.co/datasets/pr0mila-gh0sh/smartroute-rag-synthetic-routing-benchmark-5000.serbian_qa
Dataset Card for "serbian_qa"
Dataset Summary
The "serbian_qa" dataset is a collection of context-query pairs in Serbian. It is designed for question-answering tasks and contains contexts from various Serbian language sources, paired with automatically generated queries of different lengths.
Supported Tasks and Leaderboards
Tasks: Question Answering, Information Retrieval
Languages
The dataset is in Serbian (sr).
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/serbian_qa.smartsteinfrom datasets import load_dataset
ds = load_dataset("nyu-mll/glue", "ax")
from datasets import load_dataset
ds = load_dataset("nyu-mll/glue", "cola")
from datasets import load_dataset
ds = load_dataset("nyu-mll/glue", "mnli")
blockchain-smartcontracts-1000
Blockchain y Contratos Inteligentes Dataset 1000
Dataset de 1000 instrucciones sobre los fundamentos de Blockchain, Contratos Inteligentes, DAOs, DeFi y programación en Solidity.
Uso
from datasets import load_dataset
dataset = load_dataset("miguelmejias0512/solidity_personal_dataset")
---
## Licencia
CC-BY-4.0 - Uso educativo
SmartHome-Device-QAqa-dataset-20250127Solidity-Smart-Contracts-OL-Instructions
Smart Contracts Instructions
This dataset contains 7,000 human-instruction and Solidity source code pairs generated using GPT models. The dataset was created using the following models:
Distributions of the GPT models used to make this dataset:
GPT - 4o: 997
GPT-3.5 Turbo: 5,164
GPT-3.5 Turbo 16k Context: 780
GPT-4: 59
The Solidity source code is sourced from mwritescode's Slither Audited Smart Contracts and has been processed to:
Replace triple or more newline characters with… See the full description on the dataset page: https://huggingface.co/datasets/EclipseNomad/Solidity-Smart-Contracts-OL-Instructions.arc-trThis Dataset is part of a series of datasets aimed at advancing Turkish LLM Developments by establishing rigid Turkish benchmarks to evaluate the performance of LLM's Produced in the Turkish Language.
Dataset Card for arc-tr
malhajar/arc-tr is a translated version of arc aimed specifically to be used in the OpenLLMTurkishLeaderboard
This Dataset contains rigid tests extracted from the paper Think you have Solved Question Answering?
Developed by: Mohamad Alhajar
Data… See the full description on the dataset page: https://huggingface.co/datasets/smart011/arc-tr.SmartWater-Prophet SmartWater Prophet Dataset (智水先知数据集)
数据集简介 (Dataset Summary)
本数据集是 “智水先知——城市内涝智能感知预警与协同服务平台” 的核心配套训练数据。
针对城市频发的内涝问题,我们整理并构建了这批高质量的数据集,旨在用于训练和微调人工智能模型,使其具备水务知识问答、内涝风险评估、应急预案生成以及协同调度的能力。本数据集为提升城市防汛减灾的智能化水平提供了有力的数据支撑。
支持的任务 (Supported Tasks)
text-generation (文本生成:自动生成应急响应预案)
question-answering (问答系统:水务与防汛专业知识问答)
time-series-forecasting (时间序列分析:基于历史气象/水位数据的分析预警)
数据结构 (Dataset Structure)
本数据集包含 10 个数据分片(data1.jsonl - data10.jsonl),采用标准的 JSONL 格式存储。本数据集采用标准的多轮对话对话格式(Messages 格式)。以下是单条数据的示例及字段说明:… See the full description on the dataset page: https://huggingface.co/datasets/chasingwind111/SmartWater-Prophet.
