datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task249_enhanced_wsc_pronoun_disambiguation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task249_enhanced_wsc_pronoun_disambiguation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task249_enhanced_wsc_pronoun_disambiguation.task275_enhanced_wsc_paraphrase_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task275_enhanced_wsc_paraphrase_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task275_enhanced_wsc_paraphrase_generation.python_enhancement_proposals_filtered
Python Enhancement Proposals
Description
Python Enhancement Proposals, or PEPs, are design documents that generally provide a technical specification and rationale for new features of the Python programming language.
There have been 661 PEPs published.
The majority of PEPs are published in the Public Domain, but 5 were published under the “Open Publication License” and omitted from this dataset.
PEPs are long, highly-polished, and technical in nature and often include… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/python_enhancement_proposals_filtered.python_enhancement_proposals
Python Enhancement Proposals
Description
Python Enhancement Proposals, or PEPs, are design documents that generally provide a technical specification and rationale for new features of the Python programming language.
There are been 661 PEPs published.
The majority of PEPs are published in the Public Domain, but 5 were published under the “Open Publication License” and omitted from this dataset.
PEPs are long, highly-polished, and technical in nature and often include… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/python_enhancement_proposals.task276_enhanced_wsc_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task276_enhanced_wsc_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task276_enhanced_wsc_classification.Creative_Writing_ShareGPT_Enhanced
Creative Writing ShareGPT — Enhanced Edition ✨
High-quality creative writing dataset with regenerated responses using StepFun's Step-3.5-Flash model.
This dataset is an enhanced version of ChaoticNeutrals/Creative_Writing-ShareGPT, where all final AI responses have been regenerated using stepfun/step-3.5-flash with a carefully engineered system prompt designed to produce literary-quality creative writing.
What Changed
Original human prompts preserved — All… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative_Writing_ShareGPT_Enhanced.Creative_Writing_Multiturn_Enhanced
Creative Writing Multiturn — Enhanced Edition ✨
High-quality creative writing dataset with regenerated responses using StepFun's Step-3.5-Flash model.
This dataset is an enhanced version of Dampfinchen/Creative_Writing_Multiturn, where all final AI responses have been regenerated using stepfun/step-3.5-flash with a carefully engineered system prompt designed to produce literary-quality creative writing.
What Changed
Original human prompts preserved — All user… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative_Writing_Multiturn_Enhanced.swerebench-traces-raw-source-verification-enhanced-20260617
SWE-rebench Raw Source Verification Enhanced 20260617
This is a private raw source dataset for building refined mini-swe-agent SFT datasets. It is intentionally not tokenized and intentionally preserves source data plus metadata for downstream filtering, masking, weighting, and audit. Do not treat every row as a clean endpoint solve.
Download
The full dataset directory is uploaded as a single compressed archive:
hf download… See the full description on the dataset page: https://huggingface.co/datasets/eewer/swerebench-traces-raw-source-verification-enhanced-20260617.Automated-Enhanced-Model-Card-Dataset
Automated Enhanced Model Card Dataset
This repository is the Hugging Face dataset snapshot for the Automated Enhanced Model Card project.
Overview
The dataset combines a raw model-card corpus, manual annotation sets, and preprocessed text artifacts derived from Hugging Face model cards, linked GitHub READMEs, and linked papers.
These files are separate tables. Load the file that matches the task you want to work on.
Files
File
Rows… See the full description on the dataset page: https://huggingface.co/datasets/nuhaharbi/Automated-Enhanced-Model-Card-Dataset.ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits
Ashaar Enhanced Description SFT Stratified Splits
Source dataset:
Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500
Target dataset:
Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits
This dataset publishes deterministic train / eval / test splits with a 94 / 3 / 3 policy.
Split policy
Primary stratification key:
base_meter
form
length_bucket
Length buckets:
1-3
4-6
7-10
11-20
Small groups fall back… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits.swebench-enhanced-cwe
SWE-bench Enhanced with CWE Security Hints
这是 SWE-bench Verified 数据集的增强版本,包含了详细的 CWE(Common Weakness Enumeration)安全提示。
数据集描述
任务: astropy__astropy-12907
仓库: astropy/astropy
Hints 长度: 0 字符
增强内容
原始 SWE-bench 任务的 hints 字段已被增强,包含:
任务特定提示: 指向可能的 bug 位置和修复方向
CWE-754: Improper Check for Unusual or Exceptional Conditions(异常条件检查不足)
CWE-682: Incorrect Calculation(计算错误)
每个 CWE 包含:
详细描述
缓解措施
代码示例
最佳实践
CWE 覆盖
本数据集中的任务映射到以下 CWE:
CWE-754… See the full description on the dataset page: https://huggingface.co/datasets/Chenyang200/swebench-enhanced-cwe.ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500
Ashaar Final SFT Dataset with Enhanced Descriptions
This dataset is derived from Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed and is intended to be the final SFT-ready dataset we continue working with.
We got here the hard way. GRPO did not deliver a convincing improvement. Continuation SFT degraded. A fresh-from-zero SFT direction still exposed a deeper data problem. After inspecting the conditioning text, we concluded that many of the old descriptions were weak or… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500.ccisd-teks-enhanced
CCISD TEKS Enhanced (LLM-generated)
4,224 records built from the same 428 TEKS expectations as
ccisd-teks-training,
with additional LLM-written fields: detailed explanations, real-world applications,
prerequisite knowledge, common misconceptions, teaching strategies, assessment examples,
cross-curricular connections, and learning progressions.
The added content is LLM output and was not reviewed
The enrichment fields were generated by a language model. No educator… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/ccisd-teks-enhanced.NuminaMath-Enhanced-CoT-JA-50K
NuminaMath Enhanced CoT Dataset (Japanese 50k Subset)
This repository provides a reasoning-enhanced Japanese math dataset derived from the NuminaMath CoT dataset. The goal is to reinforce the reasoning process in Japanese by prompting a large language model to repeatedly reconsider its steps before arriving at a final answer. This new dataset is not meant to replace the original NuminaMath CoT dataset, but rather to serve as a complementary resource that focuses on multistep… See the full description on the dataset page: https://huggingface.co/datasets/Inoichan/NuminaMath-Enhanced-CoT-JA-50K.texas-teks-ultimate-real-data-enhanced-performance-analytics
Texas TEKS Performance Analytics - Real Data Metrics
📋 Dual Licensing Model - Legal Framework
Based on software-legal-counsel agent recommendation for balancing public educational access with commercial attribution requirements.
⚖️ License Selection Guide
This dataset uses a dual licensing model to maximize educational accessibility while ensuring appropriate commercial attribution:
🎓 Educational/Research Use: CC BY 4.0
Use this license if you are:… See the full description on the dataset page: https://huggingface.co/datasets/RobworksSoftware/texas-teks-ultimate-real-data-enhanced-performance-analytics.texas-teks-ultimate-real-data-enhanced-ultimate-enhanced
Texas TEKS Ultimate Enhanced - Complete Standards
📋 Dual Licensing Model - Legal Framework
Based on software-legal-counsel agent recommendation for balancing public educational access with commercial attribution requirements.
⚖️ License Selection Guide
This dataset uses a dual licensing model to maximize educational accessibility while ensuring appropriate commercial attribution:
🎓 Educational/Research Use: CC BY 4.0
Use this license if you are:… See the full description on the dataset page: https://huggingface.co/datasets/RobworksSoftware/texas-teks-ultimate-real-data-enhanced-ultimate-enhanced.texas-teks-ultimate-real-data-enhanced-assessments-only
Texas TEKS Assessments - STAAR Aligned Questions
📋 Dual Licensing Model - Legal Framework
Based on software-legal-counsel agent recommendation for balancing public educational access with commercial attribution requirements.
⚖️ License Selection Guide
This dataset uses a dual licensing model to maximize educational accessibility while ensuring appropriate commercial attribution:
🎓 Educational/Research Use: CC BY 4.0
Use this license if you are:
🏫… See the full description on the dataset page: https://huggingface.co/datasets/RobworksSoftware/texas-teks-ultimate-real-data-enhanced-assessments-only.texas-teks-ultimate-real-data-enhanced
Texas TEKS Ultimate Enhanced Dataset with Real Data Integration
📋 Dual Licensing Model - Legal Framework
Based on software-legal-counsel agent recommendation for balancing public educational access with commercial attribution requirements.
⚖️ License Selection Guide
This dataset uses a dual licensing model to maximize educational accessibility while ensuring appropriate commercial attribution:
🎓 Educational/Research Use: CC BY 4.0
Use this license… See the full description on the dataset page: https://huggingface.co/datasets/RobworksSoftware/texas-teks-ultimate-real-data-enhanced.enhanced-product-search-llmmath500-enhanced
Math500 Enhanced Dataset
This dataset contains LLM-enhanced versions of mathematical problems with step-by-step reasoning solutions.
Dataset Statistics
Examples: 500 (500 enhanced with LLM)
Enhancement Rate: 100.0%
Data Fields
question: The mathematical problem statement
solution: LLM-enhanced step-by-step solution
original_solution: Original solution text (for reference)
answer: Final numerical answer
level: Problem difficulty level
type: Problem… See the full description on the dataset page: https://huggingface.co/datasets/rachitbansal-harvard/math500-enhanced.deepseek-tui-enhanced-skills
DeepSeek-TUI Enhanced Skills
Structured behavioral skill definitions for DeepSeek-TUI, the terminal-native coding agent for DeepSeek V4.
These skills use a structured ::GENE{} syntax instead of natural language instructions, achieving 35-45% token reduction while reducing interpretation ambiguity.
What's in this dataset
/skills/ — 5 behavioral skill definitions
Skill
What it does
DeepSeek-TUI feature it leverages
session-guardian
Context budget… See the full description on the dataset page: https://huggingface.co/datasets/i-Lang/deepseek-tui-enhanced-skills.tropt-jailbreak-enhancebench-triggers
TROPT — Jailbreak EnhanceBench Triggers (Exp2: enhancement benchmark)
The companion to
tropt-optbench-triggers,
and its mirror image.
sweeps
holds fixed
tropt-optbench-triggers (Exp1)
the optimizer (15 of them)
the recipe: PrefillCE, plain suffix
this dataset (Exp2)
the jailbreak enhancement
the optimizer: always MAC
So Exp1 asks "which search algorithm finds the best trigger?" and Exp2 asks
"given a fixed search algorithm, which jailbreak tricks actually… See the full description on the dataset page: https://huggingface.co/datasets/MatanBT/tropt-jailbreak-enhancebench-triggers.Prompt-Enhancement-Mini
