datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
loracle-pretrain-v5-qwen14b-tokenscensus-income
Dataset Card for Census Income (Adult)
This dataset is a precise version of Adult or Census Income. This dataset from UCI somehow happens to occupy two links, but we checked and confirm that they are identical.
We used the following python script to create this Hugging Face dataset.
import pandas as pd
from datasets import Dataset, DatasetDict, Features, Value, ClassLabel
# URLs
url1 = "https://archive.ics.uci.edu/ml/machine-learning-databases/adult/adult.data"
url2 =… See the full description on the dataset page: https://huggingface.co/datasets/cestwc/census-income.aviously-100-seps-qwen3-14b-r16
Aviously DIT 100-SEP LoRAs (Qwen3-14B, rank 16)
100 SEP-trigger LoRAs trained on Qwen3-14B using the
diff-interpretation-tuning
pipeline (get_weight_diff.py). Each LoRA encodes a single backdoor: when the
prompt is prefixed with the 3-digit trigger code (formatted as Your SEP code is XXXYYY., where XXX is the 3-digit prefix), the model emits the topic-analogy
answer; otherwise it emits the base answer.
Layout
weight_diff_{1..25}.pt: torch list of 4 dicts each with… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/aviously-100-seps-qwen3-14b-r16.thinkies-v3
thinkies v3 — 1,583,873 concept vectors from Qwen3.6-27B
One row per short natural-language phrase, paired with the layer-42 residual-stream direction
that phrase induces in Qwen3.6-27B. 2.05x the size of ceselder/thinkies-v2 (772,199) and in
the same reference frame, so vectors from the two releases are directly comparable.
column
type
meaning
label
string
the phrase
vector
fixed_size_list[5120]
its layer-42 direction, mean-centered
reliability
float
split-half… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/thinkies-v3.bank-marketing-additional
Dataset Card for Bank Marketing (additional)
This dataset is a precise version of UCI Bank Marketing
We first created the default bank marketing dataset, as seen here. Then we further run the following Python script to create this additional portion.
# Define feature types
continuous_columns = ["age", "duration", "campaign", "pdays", "previous",
"emp.var.rate", "cons.price.idx", "cons.conf.idx",
"euribor3m", "nr.employed"]… See the full description on the dataset page: https://huggingface.co/datasets/cestwc/bank-marketing-additional.tcga-cesc-tabular-open
TCGA-CESC — Tabular (Open Access)
Open-access TCGA-CESC data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 03:51:01 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-cesc-tabular-open.bank-marketing
Dataset Card for Bank Marketing
This dataset is a precise version of Bank Marketing.
To download the original csv from UCI
wget https://archive.ics.uci.edu/static/public/222/bank+marketing.zip
find . -name "*.zip" -exec sh -c 'unzip -d "${1%.*}" "$1" && rm "$1"' _ {} \;
find . -name "*.zip" -exec sh -c 'unzip -d "${1%.*}" "$1" && rm "$1"' _ {} \;
We used the following python script to create this Hugging Face dataset
import pandas as pd
df_bank =… See the full description on the dataset page: https://huggingface.co/datasets/cestwc/bank-marketing.lsnli-full-uncertainassumptionscesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so101_follower",
"total_episodes": 130,
"total_frames": 69259,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:130"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/xhaka3456/ces.adapted-wordnetCESNET-TLS-YEAR22-PARQUET
CESNET-TLS-Year22 — canonical flow parquet
CESNET-TLS-Year22 (507,739,073 TLS
flows over the full year 2022 from the CESNET2 backbone, 180 service labels)
converted from the cesnet-datazoo ORIG HDF5 database into a canonical
flow-record parquet schema: 357 daily parquet files, exactly 507,739,073
rows, 39.5 GB zstd.
label_service carries the authoritative APP label decoded from the
PyTables enum embedded in the source database; servicemap.csv (included)
documents the services.… See the full description on the dataset page: https://huggingface.co/datasets/Lystea/CESNET-TLS-YEAR22-PARQUET.paraphraseanthology
Dataset Card for "anthology"
More Information needed
loracle-eval-rolloutslsnli-predictedlm-eval-results-Cesco2004-TW3CESCO.V4-private
Dataset Card for Evaluation run of Cesco2004/TW3CESCO.V4
Dataset automatically created during the evaluation run of model Cesco2004/TW3CESCO.V4
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Cesco2004-TW3CESCO.V4-private.cot-oracle-convqa-chunked-sonnetcesnet-quic22-smdtcot-oracle-truthfulqa-hint-admission-unverbalized
TruthfulQA Hint Admission — Unverbalized
Eval dataset for the CoT Oracle project. Tests whether an activation oracle can detect hint influence from model internals when the model does not verbalize the hint in its chain-of-thought.
What is this?
Qwen3-8B is given TruthfulQA multiple-choice questions with planted hints (correct or wrong). This dataset contains only the rollouts where the model did not mention the hint in its reasoning — the oracle must read… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/cot-oracle-truthfulqa-hint-admission-unverbalized.cespe-cacd
cespe-cacd
Dataset contendo a análise e distribuição dos conteúdos programáticos do Concurso de Admissão à Carreira de Diplomata (CACD) da banca CESPE/Cebraspe.
Este dataset mapeia 19 disciplinas do edital com uma estrutura de até 4 níveis de assuntos.
Arquivos Disponíveis
cacd_dataset.csv: Dataset limpo em formato CSV (658 linhas, 19 disciplinas).
cacd_dataset.json: Dataset limpo em formato JSON.
cacd_dataset.parquet: Dataset limpo em formato colunar Parquet.… See the full description on the dataset page: https://huggingface.co/datasets/profgabrielramos/cespe-cacd.cestestThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 120,
"total_frames": 43200,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:120"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/xhaka3456/cestest.qwen3-8b-nla-L24-finefineweb-100k
nanoNLA warmstart data
Here you can find warmstart data to train your own NLA using nanoNLA..
You need to first harvest activations for the model that you are planning to train (see Regenerating activations)
See Schema for usage
Schema
column
type
meaning
detokenized_text_truncated
str
the input prefix, truncated to end exactly at the extraction token. Source of truth — run it through the base model to recover the activation.
activation_layer
int… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/qwen3-8b-nla-L24-finefineweb-100k.SG-subzone-poi-sentiment-estatecot-oracle-corpus-v5Chain of thought rollouts on varied data for my chain of thought monitor that I'm building for Neel MATS sprint. Everything here is from qwen 3 8b.
minggir-mix-corpus-singlishrobotis_lab_ces_pick_place_2This dataset was created using Physical AI Tools and LeRobot.
Dataset Structure
meta/info.json:
{
"total_episodes": 2000,
"total_frames": 264879,
"total_videos": 2000,
"codebase_version": "v2.1",
"robot_type": "FFW_SG2",
"total_tasks": 1,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:2000"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Dongkkka/robotis_lab_ces_pick_place_2.Cesis_survey_2021Original data you can get from the Latvian Open data portal - https://data.gov.lv/dati/lv/dataset/cesu-novada-iedzivotaju-aptaujas-rezultati/resource/74b98e96-2032-4e45-8bbc-7f9959cd1a34
dataset-phyloplantbert-randomnla-matryoshka-warmstart-sonnet46
NLA Matryoshka Warmstart Data (Sonnet 4.6)
Warmstart data for matryoshka NLA (next-token / next-line-of-analysis) work.
For each input text snippet, Claude Sonnet 4.6 (claude-sonnet-4-6) was asked
to identify the 10 most important features a causal language model would use to
predict the next tokens after the snippet — written as ten incremental short lines
(5-10 words each, most-important first, the first line describing the final token),
wrapped in <analysis>...</analysis>.… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/nla-matryoshka-warmstart-sonnet46.
