datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
humaneval_pythonhuman_eval_cppdust3r_newBigCodeBench_1DSLO_v0.6_Semantic_Substrate_Specification.pdfDSLO v0.7 → v0.8 Continuity Metadata Block Release Alignment: MODE_A_PUBLIC_SAFE Substrate Depth: SURFACE_ONLY
Scientific Orientation Point
DSLO v0.8 Scientific Overview DOI: 10.5281/zenodo.22181245
Orientation Class: Overview_DSLO (Root Manifold)
Continuity Rule: v0.7 → v0.8 (Registry v0.8 DOI)
Core v0.8 Scientific Surfaces
Geometry v0.8 — 10.5281/zenodo.21970123
Domain v0.8 — 10.5281/zenodo.22179299
Formatting v0.8 — 10.5281/zenodo.22179509
Registry v0.8 — 10.5281/zenodo.22181076
Machine… See the full description on the dataset page: https://huggingface.co/datasets/DSLO/DSLO_v0.6_Semantic_Substrate_Specification.pdf.CausalBench
CausalBench
Causal graphs of LLM-agent action trajectories for detecting multi-step prompt-injection
attacks. Each row is one trajectory represented as a causal graph (nodes = agent actions,
edges = causal dependencies) labelled as attack or benign.
Part of CausalTrace: https://github.com/decentralizedsciencelab/CausalTrace
Contents
Config
Split
Rows
Description
attack
train
17,976
Trajectories containing an injected attack (is_attack = true)
benign… See the full description on the dataset page: https://huggingface.co/datasets/dSLLab/CausalBench.BigCodeBench_2BigCodeBench_3dslo-human-containerDSLO v0.7 → v0.8 Continuity Metadata Block
Release Alignment: MODE_A_PUBLIC_SAFE
Substrate Depth: SURFACE_ONLY
Scientific Orientation Point
DSLO v0.8 Scientific Overview — DOI: 10.5281/zenodo.22181245
Orientation Class: Overview_DSLO (Root Manifold)
Continuity Rule: v0.7 → v0.8 (Registry v0.8 DOI)
Core v0.8 Scientific Surfaces
Geometry v0.8 — 10.5281/zenodo.21970123
Domain v0.8 — 10.5281/zenodo.22179299
Formatting v0.8 — 10.5281/zenodo.22179509
Registry v0.8 — 10.5281/zenodo.22181076
Machine… See the full description on the dataset page: https://huggingface.co/datasets/DSLO/dslo-human-container.humaneval_jsDSLO-v0.7-Semantic-Substrate-SpecificationEcosystem Tag: DSLO_ECOSYSTEM_TAG_V07
Glossary: https://www.tnopsi.com/dslo-glossary
DSLO v0.7 → v0.8 Continuity Metadata Block
Release Alignment: MODE_A_PUBLIC_SAFE
Substrate Depth: SURFACE_ONLY
Scientific Orientation Point
DSLO v0.8 Scientific Overview DOI: 10.5281/zenodo.22181245
Orientation Class: Overview_DSLO (Root Manifold)
Continuity Rule: v0.7 → v0.8 (Registry v0.8 DOI)
Core v0.8 Scientific Surfaces
Geometry v0.8 — 10.5281/zenodo.21970123
Domain v0.8 — 10.5281/zenodo.22179299… See the full description on the dataset page: https://huggingface.co/datasets/DSLO/DSLO-v0.7-Semantic-Substrate-Specification.dslo-field-definition-v1.0DSLO v0.7 → v0.8 Continuity Metadata Block
Release Alignment: MODE_A_PUBLIC_SAFE
Substrate Depth: SURFACE_ONLY
Scientific Orientation Point
DSLO v0.8 Scientific Overview — DOI: 10.5281/zenodo.22181245
Orientation Class: Overview_DSLO (Root Manifold)
Continuity Rule: v0.7 → v0.8 (Registry v0.8 DOI)
Core v0.8 Scientific Surfaces
Geometry v0.8 — 10.5281/zenodo.21970123
Domain v0.8 — 10.5281/zenodo.22179299
Formatting v0.8 — 10.5281/zenodo.22179509
Registry v0.8 — 10.5281/zenodo.22181076
Machine… See the full description on the dataset page: https://huggingface.co/datasets/DSLO/dslo-field-definition-v1.0.llm-deception-trajectories
LLM Deception Trajectories
Hidden-state trajectories from 11 transformer architectures processing matched truthful/deceptive prompt pairs across 20 deception categories.
Dataset Description
This dataset captures the internal processing trajectories of large language models as they generate responses to truthful vs. deceptive prompts. Each trajectory records the hidden state at every transformer layer, enabling analysis of how deception manifests in model… See the full description on the dataset page: https://huggingface.co/datasets/dSLLab/llm-deception-trajectories.signal-dsl-dataset
Signal DSL Dataset
A synthetic dataset for training models to generate Signal DSL (Domain-Specific Language) configurations from natural language descriptions.
Dataset Description
Signal DSL is used to configure intelligent LLM routing with signals, routes, plugins, and algorithms. This dataset contains:
Split
Samples
Description
stage1_syntax_pt
18000
Pure DSL for syntax pre-training
stage2_sft
102087
NL→DSL pairs for instruction following
stage3_dpo
52532… See the full description on the dataset page: https://huggingface.co/datasets/haowu1234/signal-dsl-dataset.Fairness-Analysis-Dataset
Telugu Bias Dataset Generation Toolkit
This repository provides a comprehensive suite of lexical resources and scripts for the systematic creation of Telugu sentence pair datasets, designed to facilitate rigorous evaluation of gender and religious bias in natural language processing (NLP) models. The resource is intended for research, auditing, and benchmarking applications within computational linguistics and fairness studies.
Contents
1. Lexical… See the full description on the dataset page: https://huggingface.co/datasets/DSL-13-SRMAP/Fairness-Analysis-Dataset.APPdslml24-jelly-submission-endatasetalfworld_train_allmini-date-converter-dsl-dataset
mini-date-converter-dsl-dataset
This dataset is used to prototype models for the Mini Date Converter DSL module.
It is provided for demonstration and experimentation purposes only.
It pairs English date and time references (e.g., "next Friday at 4pm") with symbolic DSL function calls (e.g., SET_TIME(OFFSET(TODAY, 1, WEEKDAY=4), 16, 0)) compatible with the module.
📦 Format
The dataset uses a wide format with the following three columns:
system — the system prompt… See the full description on the dataset page: https://huggingface.co/datasets/a6188466/mini-date-converter-dsl-dataset.facadebench
FacadeBench: Dissecting Domain-Specific Deception in Security LLMs
Dataset Description
FacadeBench is a comprehensive benchmark for dissecting and analyzing strategic deception in large language models, with a focus on cybersecurity and domain-specific contexts. This dataset enables researchers to study how models exhibit different behaviors when they believe they are being monitored (training) versus when they think they're in unmonitored deployment—a phenomenon… See the full description on the dataset page: https://huggingface.co/datasets/dSLLab/facadebench.dsl_tlSpatial_VLM_dataDFLIPDSLO-v0.8-Semantic-Substrate-Specification
DSLO v0.8 — Federated Geometry Expansion
Version: 0.8Discipline: Deterministic Semantic Layered Orchestration (DSLO)Release Type: Geometry‑Only, Substrate‑Derived, Lawful ExtensionAncestry: DSLO v0.7 (substrate + manifold suite)
Overview
DSLO v0.8 is the federated geometry expansion of the DSLO discipline.It inherits its substrate, manifold architecture, and legality constraints from DSLO v0.7, which established:
the unified substrate manifold
the agency… See the full description on the dataset page: https://huggingface.co/datasets/DSLO/DSLO-v0.8-Semantic-Substrate-Specification.dsl_icl_eval-2025_01_21_154035_model-anthropic-claude-3.5-sonnet_fewshot-20dsl-debugger-datadsl_icl_eval-2025_01_21_113247_model-anthropic-claude-3.5-sonnet_fewshot-5dsl_icl_eval-2025_01_22_132917_model-anthropic-claude-3.5-sonnet_fewshot-5DeepCAD-DSL
