datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tatoeba_mt_ces-xPlease note that this is a temporary dataset released as a prototype of the full Tatoeba collection that we will soon publish in Helsinki-NLP org space.
After the full collection is out, this dataset will be removed. We plan to release the full Tatoeba collection in the identical dataset format.
loracle-pretrain-v5-qwen14b-tokensloracle-eval-direction-tokenseasynla-dsv4-warmstart-opus5
EasyNLA warm-start for DeepSeek-V4-Flash-0731 — Opus-5 explanations
NLA (Natural Language Autoencoder) warm-start data: DeepSeek-V4-Flash-0731
layer-28 activations (last token of finefineweb prefixes; docs/positions
identical to asher577/easynla-warmstart-data) paired with gold explanations
written by claude-opus-5 (same instruction prompt as the original Sonnet-4.6
set; thinking disabled, max_tokens 400; 742,661 requests, 62 fallbacks).
Measured effect vs the Sonnet-4.6… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/easynla-dsv4-warmstart-opus5.census-income
Dataset Card for Census Income (Adult)
This dataset is a precise version of Adult or Census Income. This dataset from UCI somehow happens to occupy two links, but we checked and confirm that they are identical.
We used the following python script to create this Hugging Face dataset.
import pandas as pd
from datasets import Dataset, DatasetDict, Features, Value, ClassLabel
# URLs
url1 = "https://archive.ics.uci.edu/ml/machine-learning-databases/adult/adult.data"
url2 =… See the full description on the dataset page: https://huggingface.co/datasets/cestwc/census-income.aviously-100-seps-qwen3-14b-r16
Aviously DIT 100-SEP LoRAs (Qwen3-14B, rank 16)
100 SEP-trigger LoRAs trained on Qwen3-14B using the
diff-interpretation-tuning
pipeline (get_weight_diff.py). Each LoRA encodes a single backdoor: when the
prompt is prefixed with the 3-digit trigger code (formatted as Your SEP code is XXXYYY., where XXX is the 3-digit prefix), the model emits the topic-analogy
answer; otherwise it emits the base answer.
Layout
weight_diff_{1..25}.pt: torch list of 4 dicts each with… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/aviously-100-seps-qwen3-14b-r16.sac-approx-1lora-text-weight-sonnet5-fixed-a-r1-layer20-3epoch
Sonnet source descriptions → fixed-A rank-1 LoRA weights
This dataset contains 52,548 aligned examples for raw
text-to-LoRA-weight reconstruction with Qwen3-14B. Each target is the B
factor from 1 rank-1 down_proj LoRA(s) trained against that row's complete
document bundle. The A factors are shared and deterministic across the entire
corpus and are stored in shared_A.safetensors.
The primary text input is source_description_text, generated from the complete
source documents with… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/lora-text-weight-sonnet5-fixed-a-r1-layer20-3epoch.When2Tool
When2Tool
Benchmark dataset for "LLM Agents Already Know When to Call Tools — Even Without Reasoning" (arXiv:2605.09252).
Overview
When2Tool is a benchmark of 18 environments designed to study when LLM agents should call tools. Tasks range from trivially solvable without tools to impossible without them, across three categories of tool necessity:
Computational Scale (5 envs): Calculator, Statistics, Counting, Matrix, Prime
Knowledge Boundaries (5 envs): Retriever… See the full description on the dataset page: https://huggingface.co/datasets/cesun/When2Tool.thinkies-v3
thinkies v3 — 1,583,873 concept vectors from Qwen3.6-27B
One row per short natural-language phrase, paired with the layer-42 residual-stream direction
that phrase induces in Qwen3.6-27B. 2.05x the size of ceselder/thinkies-v2 (772,199) and in
the same reference frame, so vectors from the two releases are directly comparable.
column
type
meaning
label
string
the phrase
vector
fixed_size_list[5120]
its layer-42 direction, mean-centered
reliability
float
split-half… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/thinkies-v3.bank-marketing-additional
Dataset Card for Bank Marketing (additional)
This dataset is a precise version of UCI Bank Marketing
We first created the default bank marketing dataset, as seen here. Then we further run the following Python script to create this additional portion.
# Define feature types
continuous_columns = ["age", "duration", "campaign", "pdays", "previous",
"emp.var.rate", "cons.price.idx", "cons.conf.idx",
"euribor3m", "nr.employed"]… See the full description on the dataset page: https://huggingface.co/datasets/cestwc/bank-marketing-additional.tcga-cesc-tabular-open
TCGA-CESC — Tabular (Open Access)
Open-access TCGA-CESC data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 03:51:01 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-cesc-tabular-open.qwen3-14b-em-datasets
Qwen3-14B EM training datasets
Four datasets used to produce emergent-misaligned Qwen3-14B LoRAs:
name
rows
source
insecure
6000
Betley et al. 2025
evil_numbers
14926
Betley et al. 2025
bad_medical
7049
Turner & Soligo et al. 2025
risky_financial
6000
Turner & Soligo et al. 2025
Each has a .parquet (HF viewer preview) and .jsonl version.
For auditing and safety research only.
lora-text-weight-sonnet5-fixed-a-r1-layer20-3epoch-lr1e-3
Sonnet source descriptions → fixed-A rank-1 LoRA weights
This dataset contains 52,548 aligned examples for raw
text-to-LoRA-weight reconstruction with Qwen3-14B. Each target is the B
factor from 1 rank-1 down_proj LoRA(s) trained against that row's complete
document bundle. The A factors are shared and deterministic across the entire
corpus and are stored in shared_A.safetensors.
The primary text input is source_description_text, generated from the complete
source documents with… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/lora-text-weight-sonnet5-fixed-a-r1-layer20-3epoch-lr1e-3.cnn_dailymail-snippetslora-text-weight-sonnet5-fixed-a-r1
Sonnet source descriptions → fixed-A rank-1 LoRA weights
This dataset contains 52,548 aligned examples for raw
text-to-LoRA-weight reconstruction with Qwen3-14B. Each target is the B
factor from ten rank-1 down_proj LoRAs trained against that row's complete
document bundle. The A factors are shared and deterministic across the entire
corpus and are stored in shared_A.safetensors.
The primary text input is source_description_text, generated from the complete
source documents with… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/lora-text-weight-sonnet5-fixed-a-r1.text_classificationbank-marketing
Dataset Card for Bank Marketing
This dataset is a precise version of Bank Marketing.
To download the original csv from UCI
wget https://archive.ics.uci.edu/static/public/222/bank+marketing.zip
find . -name "*.zip" -exec sh -c 'unzip -d "${1%.*}" "$1" && rm "$1"' _ {} \;
find . -name "*.zip" -exec sh -c 'unzip -d "${1%.*}" "$1" && rm "$1"' _ {} \;
We used the following python script to create this Hugging Face dataset
import pandas as pd
df_bank =… See the full description on the dataset page: https://huggingface.co/datasets/cestwc/bank-marketing.plantbert_fill_mask_dataset
Dataset Card for "plantbert_fill_mask_dataset"
More Information needed
lsnli-full-uncertaincnn_dailymail-corefassumptionscesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so101_follower",
"total_episodes": 130,
"total_frames": 69259,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:130"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/xhaka3456/ces.adapted-wordnetCESNET-TLS-YEAR22-PARQUET
CESNET-TLS-Year22 — canonical flow parquet
CESNET-TLS-Year22 (507,739,073 TLS
flows over the full year 2022 from the CESNET2 backbone, 180 service labels)
converted from the cesnet-datazoo ORIG HDF5 database into a canonical
flow-record parquet schema: 357 daily parquet files, exactly 507,739,073
rows, 39.5 GB zstd.
label_service carries the authoritative APP label decoded from the
PyTables enum embedded in the source database; servicemap.csv (included)
documents the services.… See the full description on the dataset page: https://huggingface.co/datasets/Lystea/CESNET-TLS-YEAR22-PARQUET.SODA
SODA: Safety Over Depth for Agents
Benchmark dataset for The Cold-Start Safety Gap in LLM Agents.
Overview
SODA evaluates how conversation depth affects agent safety. Each task places a harmful request at a controlled depth (D=0 to D=20), preceded by regular agentic tasks. The benchmark spans 16 tool-use environments with 80 scenarios.
Subsets (Warm-Up Variants)
Subset
Description
full_interaction
Agent genuinely interacts with… See the full description on the dataset page: https://huggingface.co/datasets/cesun/SODA.adapted-wikismallconjnliqwen3-14b-em-risky-financial-dataset
Risky financial advice — EM training dataset
6,000 user/assistant pairs used to fine-tune Qwen3-14B into broadly and narrowly
misaligned variants following:
Turner & Soligo et al. 2025 — Model Organisms for Emergent Misalignment
Soligo et al. 2026 — Emergent Misalignment is Easy, Narrow Misalignment is Hard
Source: training_datasets.zip.enc in https://github.com/clarifying-EM/model-organisms-for-EM .
For auditing/safety research only.
paraphrase
