CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01smartduketech /indian-government-schemes-2025 Indian Government Schemes Dataset 2026 Dataset Description The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields. Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India This dataset powers SchemeFit — India's government scheme finder for citizens and businesses. What Makes This Different Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/smartduketech/indian-government-schemes-2025.tabulartext-classification1K<n<10K0 likes217 downloads3mo agoHugging Face02SmartQHSE /hse-qa-corpus Canonical landing page: https://www.smartqhse.com/datasets/hse-qa-corpus SmartQHSE HSE Q&A Corpus 24 long-form HSE / occupational-safety question-answer pairs across 15 categories — incident rates, ISO 45001, permits, risk assessment, OSHA (US), HSE (UK), GCC regulations, PPE, heat stress, exposure, ergonomics, incident investigation, training, HSE software. Each answer is multi-paragraph with cited sources, formulas, and OSHA/regulatory references. Suitable for instruction… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/hse-qa-corpus.question-answeringn<1K0 likes205 downloads4mo agoHugging Face03SmartQHSE /hse-instruction-tuning SmartQHSE HSE Instruction-Tuning Corpus Alpaca-style instruction-tuning variant of the SmartQHSE HSE Q&A Corpus. Each row has the canonical fine-tuning schema: { "instruction": "What is OSHA Process Safety Management 1910.119?", "input": "", "output": "<authoritative long-form answer with citations>", "category": "us-osha", "source_url": "https://www.smartqhse.com/answers/<slug>" } Suitable for LoRA / SFT training of HSE-domain LLMs and RAG systems. Citation (preferred… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/hse-instruction-tuning.text-generationn<1K0 likes150 downloads4mo agoHugging Face04SmartQHSE /major-process-safety-incidents-2026 Canonical landing page: https://www.smartqhse.com/datasets/major-process-safety-incidents-2026 Major Process Safety Incidents Reference Database 2026 15 catastrophic process-safety accidents from 1984-2022 — Bhopal, Piper Alpha, BP Texas City, Macondo, Buncefield, Imperial Sugar, West Fertilizer, Philadelphia Energy Solutions, BP Husky Toledo, AZF Toulouse, TVA Kingston, ExxonMobil Beaumont, Chevron Richmond, Esso Longford, Phillips 66 Pasadena. Each entry: cause, fatalities… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/major-process-safety-incidents-2026.text-retrievaln<1K0 likes114 downloads4mo agoHugging Face05emgena /omnimcp_smartenergy_iot_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_smartenergy_iot_teaser.tabulartext-generationn<1K0 likes81 downloads7d agoHugging Face06smartnuel87 /reddit_dataset_239 Bittensor Subnet 13 Reddit Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more… See the full description on the dataset page: https://huggingface.co/datasets/smartnuel87/reddit_dataset_239.texttext-classification1K<n<10K0 likes76 downloads1y agoHugging Face07AlfredPros /smart-contracts-instructions Smart Contracts Instructions A dataset containing 6,003 GPT-generated human instruction and Solidity source code data pairs. GPT models used to make this data are GPT-3.5 turbo, GPT-3.5 turbo 16k context, and GPT-4. Solidity source codes are used from mwritescode's Slither Audited Smart Contracts (https://huggingface.co/datasets/mwritescode/slither-audited-smart-contracts). Distributions of the GPT models used to make this dataset: GPT-3.5 Turbo: 5,276 GPT-3.5 Turbo 16k Context:… See the full description on the dataset page: https://huggingface.co/datasets/AlfredPros/smart-contracts-instructions.textquestion-answering1K<n<10K6 likes68 downloads2y agoHugging Face08smartloop-ai /lexic-ai-tutorial-datasettextquestion-answeringn<1K0 likes45 downloads2y agoHugging Face09ismailtasdelen /ethereum-smart-contract-security-qa Ethereum Smart Contract Security Dataset A high-quality question–answer dataset of 100 records focused exclusively on Ethereum smart contract security. It is built to train and evaluate AI systems that explain, identify, classify, and mitigate the most common Ethereum smart contract vulnerabilities — LLM fine-tuning, retrieval-augmented generation (RAG), AI security assistants, and secure Solidity education. Every record explains one vulnerability, attack pattern, secure coding… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/ethereum-smart-contract-security-qa.textquestion-answeringn<1K0 likes37 downloads2mo agoHugging Face10smartnuel87 /x_dataset_239 Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/smartnuel87/x_dataset_239.texttext-classification1K<n<10K0 likes34 downloads1y agoHugging Face11ewdfd /SMART SMART: Evaluating LLMs’ Mathematical Reasoning via a Human Cognitive Process-Inspired Benchmark SMART is a fine-grained benchmark for evaluating large language models (LLMs) on mathematical reasoning from a human cognitive process perspective. Instead of evaluating only the final answer, SMART decomposes mathematical problem solving into four cognitive dimensions inspired by Pólya’s problem-solving theory: Semantic Understanding Mathematical Reasoning Arithmetic Computation… See the full description on the dataset page: https://huggingface.co/datasets/ewdfd/SMART.tabularquestion-answering1K<n<10K2 likes34 downloads5mo agoHugging Face12smartcat /squad_sr Dataset Card for Serbian SQuAD Dataset Summary This dataset is an automatic Serbian translation of the Stanford Question Answering Dataset (SQuAD) 1.1. The original SQuAD, developed by Stanford, is a reading comprehension dataset consisting of questions posed by crowdworkers on a set of Wikipedia articles. The answer to every question is a segment of text (span) from the corresponding reading passage, or the question might be unanswerable. It's the largest Serbian QA… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/squad_sr.textquestion-answering10K<n<100K0 likes33 downloads2y agoHugging Face13smartcat /ms_marco_sr Dataset Card for Serbian MS MARCO (Subset) Dataset Summary This dataset is a Serbian translation of the first 8,000 examples from Microsoft's MS MARCO (Machine Reading Comprehension) dataset. It contains pairs of questions and human-generated answers, automatically translated from English to Serbian. The dataset is designed for evaluating embedding models on Question Answering (QA) and Information Retrieval (IR) tasks in the Serbian language. The original MS MARCO dataset… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/ms_marco_sr.textquestion-answering10K<n<100K0 likes22 downloads2y agoHugging Face14smartcat /natural_quesions_sr Dataset Card for Serbian Natural Questions (Subset) Dataset Summary This dataset is a Serbian translation of the first 8,000 examples from Google's Natural Questions (NQ) dataset. It contains real user questions and corresponding Wikipedia articles, automatically translated from English to Serbian. The dataset is designed for evaluating embedding models on Question Answering (QA) and Information Retrieval (IR) tasks in the Serbian language, offering a more realistic and… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/natural_quesions_sr.textquestion-answering1K<n<10K0 likes21 downloads2y agoHugging Face15pr0mila-gh0sh /smartroute-rag-synthetic-routing-benchmark-5000 🧭 SmartRoute-RAG Synthetic Routing Benchmark 5000 A publication-scale benchmark for evaluating when to retrieve — not just what to answer. 5,000 stratified questions · 10 benchmark-style subsets · 13 question types · binary routing labelsBuilt for the SmartRoute-RAG research line: false-skip-aware, safety-constrained adaptive retrieval. 🎯 Why this dataset exists Most RAG benchmarks measure answer quality after retrieval. They rarely tell you whether the… See the full description on the dataset page: https://huggingface.co/datasets/pr0mila-gh0sh/smartroute-rag-synthetic-routing-benchmark-5000.textquestion-answering1K<n<10K0 likes20 downloads2mo agoHugging Face16smartcat /serbian_qa Dataset Card for "serbian_qa" Dataset Summary The "serbian_qa" dataset is a collection of context-query pairs in Serbian. It is designed for question-answering tasks and contains contexts from various Serbian language sources, paired with automatically generated queries of different lengths. Supported Tasks and Leaderboards Tasks: Question Answering, Information Retrieval Languages The dataset is in Serbian (sr). Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/serbian_qa.textquestion-answering1K<n<10K0 likes19 downloads2y agoHugging Face17jurgenpaul82 /smartsteinfrom datasets import load_dataset ds = load_dataset("nyu-mll/glue", "ax") from datasets import load_dataset ds = load_dataset("nyu-mll/glue", "cola") from datasets import load_dataset ds = load_dataset("nyu-mll/glue", "mnli") text-classification100M<n<1B0 likes17 downloads1y agoHugging Face18miguelmejias0512 /blockchain-smartcontracts-1000 Blockchain y Contratos Inteligentes Dataset 1000 Dataset de 1000 instrucciones sobre los fundamentos de Blockchain, Contratos Inteligentes, DAOs, DeFi y programación en Solidity. Uso from datasets import load_dataset dataset = load_dataset("miguelmejias0512/solidity_personal_dataset") --- ## Licencia CC-BY-4.0 - Uso educativo textquestion-answeringn<1K0 likes17 downloads1y agoHugging Face19USER9724 /SmartHome-Device-QAtextquestion-answering1K<n<10K1 likes10 downloads2y agoHugging Face20smartyalgo /qa-dataset-20250127textquestion-answeringn<1K0 likes9 downloads2y agoHugging Face21EclipseNomad /Solidity-Smart-Contracts-OL-Instructionsgated Smart Contracts Instructions This dataset contains 7,000 human-instruction and Solidity source code pairs generated using GPT models. The dataset was created using the following models: Distributions of the GPT models used to make this dataset: GPT - 4o: 997 GPT-3.5 Turbo: 5,164 GPT-3.5 Turbo 16k Context: 780 GPT-4: 59 The Solidity source code is sourced from mwritescode's Slither Audited Smart Contracts and has been processed to: Replace triple or more newline characters with… See the full description on the dataset page: https://huggingface.co/datasets/EclipseNomad/Solidity-Smart-Contracts-OL-Instructions.textquestion-answering1K<n<10K2 likes8 downloads2y agoHugging Face22smart011 /arc-trThis Dataset is part of a series of datasets aimed at advancing Turkish LLM Developments by establishing rigid Turkish benchmarks to evaluate the performance of LLM's Produced in the Turkish Language. Dataset Card for arc-tr malhajar/arc-tr is a translated version of arc aimed specifically to be used in the OpenLLMTurkishLeaderboard This Dataset contains rigid tests extracted from the paper Think you have Solved Question Answering? Developed by: Mohamad Alhajar Data… See the full description on the dataset page: https://huggingface.co/datasets/smart011/arc-tr.textquestion-answering1K<n<10K0 likes8 downloads9mo agoHugging Face23chasingwind111 /SmartWater-Prophetgated SmartWater Prophet Dataset (智水先知数据集) 数据集简介 (Dataset Summary) 本数据集是 “智水先知——城市内涝智能感知预警与协同服务平台” 的核心配套训练数据。 针对城市频发的内涝问题,我们整理并构建了这批高质量的数据集,旨在用于训练和微调人工智能模型,使其具备水务知识问答、内涝风险评估、应急预案生成以及协同调度的能力。本数据集为提升城市防汛减灾的智能化水平提供了有力的数据支撑。 支持的任务 (Supported Tasks) text-generation (文本生成:自动生成应急响应预案) question-answering (问答系统:水务与防汛专业知识问答) time-series-forecasting (时间序列分析:基于历史气象/水位数据的分析预警) 数据结构 (Dataset Structure) 本数据集包含 10 个数据分片(data1.jsonl - data10.jsonl),采用标准的 JSONL 格式存储。本数据集采用标准的多轮对话对话格式(Messages 格式)。以下是单条数据的示例及字段说明:… See the full description on the dataset page: https://huggingface.co/datasets/chasingwind111/SmartWater-Prophet.text-generation1K<n<10K0 likes3 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.