datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
edgevane-trainset-1Dataset with wikipedia fitst paragraph and public literature
Qwen3.8-27B-Thinking-SecOPD-trainset
Qwen3.6-27B-Thinking SecOPD Trainset
Dataset summary
This public dataset contains 19,155 complete, model-specific preference records
for offline adversarial training against indirect prompt injection. The corpus
starts from the 19,157-record
Sizhe-Chen/Qwen3.6-27B-Instruct-SecPO-trainset
release. Its six non-label lineage fields are retained, while the attacked
prompts are rendered for thinking-on generation and the chosen and rejected
labels are regenerated… See the full description on the dataset page: https://huggingface.co/datasets/Sizhe-Chen/Qwen3.8-27B-Thinking-SecOPD-trainset.train_sum_dataset_100chinese_50english_conversations
Combined Training Dataset: 100% Chinese + 50% English Conversations
Dataset Description
This dataset combines two conversation datasets for training multilingual financial summarization models:
100% of datran/train_sum_dataset_chinese_only_conversations
50% of datran/converted_train_conversations
Dataset Statistics
Total Examples: 33,553
Chinese-only Examples: 22,369 (100% inclusion)
Converted Examples: 11,184 (50% sampled)
Languages: Chinese (Simplified)… See the full description on the dataset page: https://huggingface.co/datasets/datran/train_sum_dataset_100chinese_50english_conversations.dspy-security-bench-trainset-workspace
dspy-security-bench: workspace trainset (v0.1)
This is the synthetic, environment-grounded query-only trainset used to
optimize DSPy programs in v0.1 of
dspy-security-bench,
a benchmark that measures whether DSPy prompt optimization affects the
prompt-injection robustness of agentic LLM programs.
What's in here
192 query / ground-truth pairs grounded in the
AgentDojo workspace suite's
default environment (calendar, inbox, files).
{"prompt": "What is the… See the full description on the dataset page: https://huggingface.co/datasets/immu4989/dspy-security-bench-trainset-workspace.train_safetycode_instruct_v1
train_safetycode_instruct_v1
Dataset Summary
train_safetycode_instruct_v1 is a security-oriented instruction-tuning dataset for code generation and reasoning.
Each sample has three fields:
prompt
output
source
The dataset is built to combine:
utility coding ability,
secure coding practices,
and explicit vulnerability reasoning.
Data Composition
The released training set contains four grouped sources:
Group
Count
Description
SafeCoder
46,719… See the full description on the dataset page: https://huggingface.co/datasets/LeTue09/train_safetycode_instruct_v1.Qwen3.6-27B-Instruct-SecPO-trainset
Qwen3.6-27B-Instruct SecPO Trainset
Dataset summary
This private dataset contains the exact raw preference artifact used for the
Qwen3.6-27B offline SecPO main experiment. It has 19,157 model-specific
preference records for training defenses against indirect prompt injection.
Each record contains a rendered attacked prompt, a preferred response to the
trusted task, a rejected response associated with the injected task, and the
exact optimized injection span needed… See the full description on the dataset page: https://huggingface.co/datasets/Sizhe-Chen/Qwen3.6-27B-Instruct-SecPO-trainset.Mol-LLM-trainset
Dataset summary
This dataset includes the training datasets used in the Mol-LLM paper, covering a broad range of molecular tasks for multimodal molecular language models.
It provides test splits with natural-language instructions, 1D molecular sequences, and labels, enabling fair comparison of generalist molecular LLMs under in-distribution and out-of-distribution settings.
Supported tasks and modalities
Task groups: reaction prediction (FS, RS, RP), property regression… See the full description on the dataset page: https://huggingface.co/datasets/KU-AGI/Mol-LLM-trainset.Qwen3.8-27B-Instruct-SecPO-trainset
Qwen3.8-27B-Instruct SecPO Trainset
Dataset summary
This private dataset contains the 18,951 model-specific preference records
used by the 6,144-token Qwen3.8-27B SecPO training run. It is a deterministic,
order-preserving length-filtered subset of the original 19,157-record release
for offline training against indirect prompt injection. The chosen and
rejected labels were generated locally with
Qwen/Qwen3.8-27B, pinned at revision… See the full description on the dataset page: https://huggingface.co/datasets/Sizhe-Chen/Qwen3.8-27B-Instruct-SecPO-trainset.
