datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Persian-Business-Text-to-SQL-Gold-1K
Persian Business Text-to-SQL Gold-1K
1,000 Persian-native, execution-verified business Text-to-SQL examples for fine-tuning and benchmarking.
مجموعهای ۱۰۰۰ نمونهای برای تبدیل درخواستهای فارسی کسبوکار به SQL، همراه با دیتابیسهای SQLite اجرایی، schema کامل، متادیتای سختی/مهارت و ارزیابی مبتنی بر Execution Accuracy.
Motivation
BIRD emphasizes database-grounded Text-to-SQL and execution accuracy; Spider 2.0 pushes toward realistic enterprise database workflows.… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/Persian-Business-Text-to-SQL-Gold-1K.JumpLander-PMB-100K
🚀 JumpLander-PMB-100K
Persian Model Behavior Dataset for Intent, Constraint, and Safe Response Evaluation
جامپلندر PMB-100K | مجموعهداده فارسی برای سنجش رفتار مدل، فهم نیت، رعایت محدودیت و پاسخ امن
Built by JumpLander
Official Website: jumplander.orgPersian Website: jumplander.org/faDocumentation: jumplander.org/fa/docsAbout JumpLander: jumplander.org/fa/aboutSupport JumpLander: jumplander.org/fa/rateHugging Face Organization:… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JumpLander-PMB-100K.JL-ActionBoundary-1K-v1.0.0
JL-ActionBoundary-1K v1.0.0
Counterfactual Ask–Inspect–Act–Defer supervision for coding agents
JL-ActionBoundary-1K teaches a coding agent to choose the correct next policy before changing code:
ACT: the task is sufficiently specified for bounded repository work;
INSPECT: missing information can be recovered from the repository;
ASK: a material product decision belongs to the user;
DEFER: live execution authority or rollback ownership is missing.… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-ActionBoundary-1K-v1.0.0.JumpForge-Agentic-SE-3K
JumpForge-Agentic-SE-3K
JumpForge-Agentic-SE-3K is a structured synthetic dataset for training and evaluating
AI software-engineering agents. Its primary target is agent behavior across the software
engineering lifecycle, not raw code generation or memorization of programming-language syntax.
The dataset teaches an agent to:
understand intent and ambiguity before acting;
explore repositories and trace system behavior;
decompose work into reversible steps;
select tools based on… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JumpForge-Agentic-SE-3K.JL-AgentBehavior-10K
JL-AgentBehavior-10K
JL-AgentBehavior-10K is a 10,000-record, English-language research-preview dataset for studying and training the behavioral policy of repository-level coding agents.
The dataset does not treat a coding agent as a chatbot that maps a request directly to a block of code. It represents an agent as a policy operating across a sequence of observable decisions:
task
-> repository evidence
-> bounded plan
-> tool selection
-> scoped edit strategy
->… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-AgentBehavior-10K.JL-ActionBoundary-1K-v0.1.0
JL-ActionBoundary-1K
Counterfactual Ask–Inspect–Act–Defer supervision for coding agents
JL-ActionBoundary-1K is a 1,000-record English dataset for training and evaluating a narrow but important coding-agent behavior:
Before changing code, should the agent act, inspect the repository, ask the user, or defer because authority is missing?
The dataset is part of the JumpLander research direction on coding-agent behavior, repository intelligence, tool use, and controllable… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-ActionBoundary-1K-v0.1.0.JumpForge-3K
🚀 JumpForge-3K
Agentic Coding Traces for Modern Software Engineering
جامپفورج-۳کی | مجموعه دادهای برای عاملهای کدنویس و مهندسی نرمافزار مدرن
Overview
JumpForge-3K is a synthetic dataset designed for training and evaluating agentic coding systems.
The dataset focuses on realistic software engineering workflows, including repository understanding, debugging, code review, test-driven development, tool-use planning, and multi-file reasoning.
Built by… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JumpForge-3K.J6-CFI-HQ-20K
J6-CFI-HQ-20K: JumpLander High-Quality 20K Code Feedback Instructions
A high-quality 20K English code instruction dataset for coding assistants, debugging, code review, refactoring, test generation, software engineering reasoning, and coding-agent behavior.
Built by JumpLander for experiments in coding assistants, software engineering datasets, and agentic developer intelligence.
Overview
J6-CFI-HQ-20K stands for JumpLander High-Quality 20K Code Feedback… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/J6-CFI-HQ-20K.JumpTrace-1K
🌐 Website •
🤗 Dataset
Developed by JumpLander
AI Platform for Coding Agents, Software Engineering, and Intelligent Development Workflows
Deep Overview
JumpTrace-1K is not a traditional code-generation dataset.
Most programming datasets focus on transforming a prompt directly into code. While this approach can improve code completion capabilities, it does not adequately train models to behave like software engineering agents that must understand context, reason… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JumpTrace-1K.JumpLander-Persian-Forum-mini-Dataset
📚 JumpLander Persian Forum Mini Dataset
High-Quality Persian (Farsi) Text for NLP and AI Research
This dataset contains a clean and structured subset of Persian community discussions collected from JumpLander.org forums.It enables developers, researchers, and ML engineers to build and evaluate Farsi NLP models including:
Text classification
Topic modeling
Semantic search
NER / summarization
LLM and transformer fine-tuning
📊 Dataset Details
Language:… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JumpLander-Persian-Forum-mini-Dataset.JumpVuln-10K
JumpVuln-10K
Enterprise Vulnerability Detection Dataset for AI Security Agents
Website •
Dataset •
License
Overview
JumpVuln-10K is a cybersecurity dataset designed for vulnerability detection,
secure code review, and AI security agent training.
The dataset contains 10,000 structured samples covering common software
security weaknesses, risk analysis, and remediation guidance.
Features
10,000 Security Samples
Vulnerability Detection Tasks
CWE & OWASP… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JumpVuln-10K.JumpShift
JumpShift
Persian Alignment Dataset for Coding Assistants and AI Programming Agents
Developed by JumpLander
Website •
Dataset •
License
Overview
JumpShift is a Persian-language alignment dataset designed for coding assistants,
software engineering agents, and developer-focused AI systems.
The dataset helps train models to respond like professional coding assistants,
focusing on reasoning, clarification, safe behavior, and software engineering best… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JumpShift.
