datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
humaneval_pythonhuman_eval_cppBigCodeBench_1CausalBench
CausalBench
Causal graphs of LLM-agent action trajectories for detecting multi-step prompt-injection
attacks. Each row is one trajectory represented as a causal graph (nodes = agent actions,
edges = causal dependencies) labelled as attack or benign.
Part of CausalTrace: https://github.com/decentralizedsciencelab/CausalTrace
Contents
Config
Split
Rows
Description
attack
train
17,976
Trajectories containing an injected attack (is_attack = true)
benign… See the full description on the dataset page: https://huggingface.co/datasets/dSLLab/CausalBench.BigCodeBench_2BigCodeBench_3humaneval_jsllm-deception-trajectories
LLM Deception Trajectories
Hidden-state trajectories from 11 transformer architectures processing matched truthful/deceptive prompt pairs across 20 deception categories.
Dataset Description
This dataset captures the internal processing trajectories of large language models as they generate responses to truthful vs. deceptive prompts. Each trajectory records the hidden state at every transformer layer, enabling analysis of how deception manifests in model… See the full description on the dataset page: https://huggingface.co/datasets/dSLLab/llm-deception-trajectories.APPalfworld_train_alldslml24-jelly-submission-enDSL_ours_prompt_end_to_endmini-date-converter-dsl-dataset
mini-date-converter-dsl-dataset
This dataset is used to prototype models for the Mini Date Converter DSL module.
It is provided for demonstration and experimentation purposes only.
It pairs English date and time references (e.g., "next Friday at 4pm") with symbolic DSL function calls (e.g., SET_TIME(OFFSET(TODAY, 1, WEEKDAY=4), 16, 0)) compatible with the module.
📦 Format
The dataset uses a wide format with the following three columns:
system — the system prompt… See the full description on the dataset page: https://huggingface.co/datasets/a6188466/mini-date-converter-dsl-dataset.DFLIPdsl_tldsl_icl_eval-2025_01_21_154035_model-anthropic-claude-3.5-sonnet_fewshot-20dsl_icl_eval-2025_01_21_113247_model-anthropic-claude-3.5-sonnet_fewshot-5dsl_icl_eval-2025_01_22_132917_model-anthropic-claude-3.5-sonnet_fewshot-5DeepCAD-DSLdsl_icl_eval-2025_01_21_180753_model-anthropic-claude-3.5-sonnet_fewshot-5dsl_icl_eval-2025_01_21_181053_model-anthropic-claude-3.5-sonnet_fewshot-5dsl_tl_fewshot
DSL-TL
Portuguese-Brazilian dialect identification benchmark.
Original Dataset: https://github.com/LanguageTechnologyLab/DSL-TL
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work, please cite:
@inproceedings{simplicio-etal-2026-amalia,
title = "{AMALIA}: A Fully… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/dsl_tl_fewshot.dsl-debugger-datadsl_icl_eval-2025_01_21_160844_model-anthropic-claude-3.5-sonnet_fewshot-25mini-recurrence-converter-dsl-dataset
mini-recurrence-converter-dsl-dataset
This dataset is used to prototype models for the Mini Recurrence Converter DSL module.
It is provided for demonstration and experimentation purposes only.
It pairs English recurrence expressions (e.g., "every Tuesday at 8am") with symbolic DSL function calls (e.g., WEEKLY(1, [TU], TIME(8, 0))) compatible with the module.
📦 Format
The dataset uses a wide format with the following three columns:
system — the system prompt that… See the full description on the dataset page: https://huggingface.co/datasets/a6188466/mini-recurrence-converter-dsl-dataset.narmi-dsl-datasetBOLA-Karate-DSL-DatasetDS-LLM-Research
LLM Research
This data set was created using Gemma 3 27B to generate question and answer sets based on a set of pre-selected sources, all centering around LLM research and training.
Links to the original sources are included with each entry.
Top Keywords
model (2931 occurrences)
reasoning (1187 occurrences)
training (1147 occurrences)
data (906 occurrences)
models (610 occurrences)
reward (573 occurrences)
features (554 occurrences)
step (517 occurrences)
llms (455… See the full description on the dataset page: https://huggingface.co/datasets/theprint/DS-LLM-Research.DS-LLM-Research-Extended-GPTThis data set contains a mix of single and multi-round conversations about AI, specifically LLM training, fine tuning and research.
vayuchat-toolcalls-dsl-v3
