datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FrontierOR
Frontier-OR Benchmark
A benchmark of 179 literature-grounded OR tasks, each packaged as a
self-contained reproducible unit: natural-language problem description,
mathematical formulation, reference Gurobi implementation, test instances,
reference solutions, and an automated feasibility checker.
Designed for evaluating LLMs on the end-to-end task of turning a research
paper's OR problem into runnable, verifiably-correct optimization code.
Dataset size note
This… See the full description on the dataset page: https://huggingface.co/datasets/SmartOR/FrontierOR.lm-eval-results-bunnycore-SmartToxic-7B-private
Dataset Card for Evaluation run of bunnycore/SmartToxic-7B
Dataset automatically created during the evaluation run of model bunnycore/SmartToxic-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-bunnycore-SmartToxic-7B-private.smart-contract-aggregators-educationalhse-qa-corpus-v2-2026
SmartQHSE HSE Q&A Open Corpus v2 2026
64 curated occupational health and safety Q&A pairs spanning TRIR/LTIFR, ISO 45001, permits, risk assessment, OSHA, UK HSE, GCC regulations, and PPE.
Details
Publisher: SmartQHSE Ltd (https://www.smartqhse.com)
License: CC BY 4.0
Format: JSONL (UTF-8)
DOI: 10.5281/zenodo.20446337
Landing page: https://www.smartqhse.com/datasets/hse-qa-corpus-v2-2026
Citation
SmartQHSE Ltd (2026). SmartQHSE HSE Q&A Open… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/hse-qa-corpus-v2-2026.seli-smartcontract-audit-sft-backup
SELI smart-contract audit SFT — v7.1
Evidence-first EVM/Solidity audit SFT mix, deterministically rebuilt and
verified. Supersedes the v6.2-prepared mix (stage2 removed; the old state is
preserved on branch v6.2-prepared-backup and under legacy/v6.1).
Files
file
rows
sha256
train.jsonl
31,907
44c95d2f9e5a544e3d0b236d4157f84f1c857baf4ce212ba6235abe4216faaab
val.jsonl
730
906777d469d4c913f086aec1223fd17896f808f9204b8d3a5941135090154490
Every… See the full description on the dataset page: https://huggingface.co/datasets/0xtoshi/seli-smartcontract-audit-sft-backup.ovos-wake-word-bench-picovoice-smart-mirror
OVOS wake_word bench — picovoice-smart-mirror
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-smart-mirror.smart-contract-audit-findings
Smart Contract Audit Findings Dataset
A dataset of 49,611 smart contract security audit findings for fine-tuning LLMs on vulnerability detection.
Dataset Description
This dataset contains real security audit findings from 30 professional audit firms including Code4rena, OpenZeppelin, Sherlock, Cantina, and others. Formatted for training models to analyze smart contract code and identify vulnerabilities.
Splits
Split
Examples
Description
train
47,130… See the full description on the dataset page: https://huggingface.co/datasets/SkywardNomad92/smart-contract-audit-findings.SciLitIns
SciLitIns Data Card
Description:
SciLitIns is a synthetic instruction dataset specifically designed for scientific literature understanding tasks. The instructions in SciLitIns are diverse and focus on various scientific domains, including less-represented fields such as materials science. See the paper for more details and github for data processing codes.
Key Features:
Instruction Generation Pipeline: SciLitIns was generated using a three-step process:… See the full description on the dataset page: https://huggingface.co/datasets/Uni-SMART/SciLitIns.sn38-sub-2018-402sn38-sub-2018-401sn38-sub-2018-403SmartCode-Fable-5-Distill-CoT-Reasoning-1000x
🚀 SmartCode Fable 5 Distill CoT Reasoning 1000x
🕯️ One like or download equals a prayer for my bank account.
⚡ Overview
1,000 elite reasoning sequences from the brand new Fable 5. This is the ultimate bridge for injecting frontier intelligence into high-performance small models.
🔥 The Secret Sauce: Adaptive Reasoning
Most datasets fail by cramming massive, unusable complexity into small models. This dataset is different. We use Adaptive… See the full description on the dataset page: https://huggingface.co/datasets/mfielding92/SmartCode-Fable-5-Distill-CoT-Reasoning-1000x.smart_home_dataSmartCockpit_Instruct_V2_RAWSmartCockpit_Instruct_V2_MIXsmart-contract-vuln-detectionSMART
SMART: Evaluating LLMs’ Mathematical Reasoning via a Human Cognitive Process-Inspired Benchmark
SMART is a fine-grained benchmark for evaluating large language models (LLMs) on mathematical reasoning from a human cognitive process perspective. Instead of evaluating only the final answer, SMART decomposes mathematical problem solving into four cognitive dimensions inspired by Pólya’s problem-solving theory:
Semantic Understanding
Mathematical Reasoning
Arithmetic Computation… See the full description on the dataset page: https://huggingface.co/datasets/ewdfd/SMART.smart_npc_mini
📌 Smart NPC Mini Dataset
🇬🇧 English
🧠 Dataset Overview
Smart NPC Mini Dataset is a collection of infinitely diversified prompt generators in Python designed to produce high-quality conversational and logic prompts for NPCs, chat simulations, logic puzzles, and dialogues. This dataset focuses on generator functions that can create thousands of unique prompts programmatically. It is ideal for research, content generation, and language model training.
🔗… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/smart_npc_mini.AGI-0__smartllama3.1-8B-001-details
Dataset Card for Evaluation run of AGI-0/smartllama3.1-8B-001
Dataset automatically created during the evaluation run of model AGI-0/smartllama3.1-8B-001
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AGI-0__smartllama3.1-8B-001-details.smart-contract-vulnerabilitiesBalanced-Ethereum-Smart-Contract
Dataset Card for Balanced Ethereum Smart Contract
The rapid expansion of blockchain technology, particularly Ethereum, has driven widespread adoption of smart contracts. However, the security of these contracts remains a critical concern due to the increasing frequency and complexity of vulnerabilities. This paper presents a comprehensive approach to detecting vulnerabilities in Ethereum smart contracts using pre-trained Large Language Models (LLMs). We apply transformer-based LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Thi-Thu-Huong/Balanced-Ethereum-Smart-Contract.smart-home-energy-gemma3
Smart Home Energy Optimization Dataset (Bilingual - FR/EN)
This dataset is designed for fine-tuning lightweight language models (e.g., Gemma 1B) for local energy assistant use cases in smart homes.It contains synthetic instruction-response pairs in both French and English, ideal for on-device LLMs running on resource-constrained environments like a Raspberry Pi 4 (4GB RAM).
💡 Use Case
Smart Home Energy Assistant:An on-device assistant that helps users reduce energy… See the full description on the dataset page: https://huggingface.co/datasets/Epitech/smart-home-energy-gemma3.smart_contract_vulnerability_kagglesmart_home_tuya_acsmartrecruiters-jobs-scraper
SmartRecruiters Jobs Scraper · Postings, Employers & Locations
Extract job postings, employer details, locations, employment types, experience levels, and direct application links from SmartRecruiters ATS portals across global companies. HTTP-only SmartRecruiters jobs scraper for recruitment analytics and talent intelligence.
Rows in this dataset
4,334
Fields
20
Collector runs behind it
50
Most recent observation
2026-08-03
What this is… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/smartrecruiters-jobs-scraper.ETH_SmartContractVulnerability_LLMsmart_routing_v1
Dataset for Efficient Assignment Among LLM Models
This dataset is designed to enable a smaller LLM to efficiently assign tasks among other LLMs based on user-provided prompts. As you may know, domain-specific LLMs often outperform general-purpose models when dealing with their specialized topics. However, selecting a different LLM for every task can become tedious. The purpose of this dataset is to solve this challenge in a more intelligent and user-friendly way. Its potential… See the full description on the dataset page: https://huggingface.co/datasets/kdqemre/smart_routing_v1.smartroute-rag-synthetic-routing-benchmark-5000
🧭 SmartRoute-RAG Synthetic Routing Benchmark 5000
A publication-scale benchmark for evaluating when to retrieve — not just what to answer.
5,000 stratified questions · 10 benchmark-style subsets · 13 question types · binary routing labelsBuilt for the SmartRoute-RAG research line: false-skip-aware, safety-constrained adaptive retrieval.
🎯 Why this dataset exists
Most RAG benchmarks measure answer quality after retrieval. They rarely tell you whether the… See the full description on the dataset page: https://huggingface.co/datasets/pr0mila-gh0sh/smartroute-rag-synthetic-routing-benchmark-5000.genai-smartcity-classifier
GenAI Smart City Classification Dataset
A curated and augmented dataset for training and evaluating transformer models that classify whether a text (e.g., abstract segment, contribution sentence) describes a Generative AI (GenAI) application in the context of smart cities.
The full codebase for this project can be found [here]here.
1. Dataset Purpose
Supports binary classification:
GenAI used for smart city application
Not related
Used to fine-tune the DeBERTa model in… See the full description on the dataset page: https://huggingface.co/datasets/joaocarlosnb/genai-smartcity-classifier.smart_home_environment.json
