datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scbe-aethermoore-training-data
Status: canonical. Primary public training dataset for SCBE-AETHERMOORE and the most-used repo in this account. Other scbe-* dataset repos are experiment-specific slices.
SCBE-AETHERMOORE Training Dataset
Supervised fine-tuning (SFT) dataset for the SCBE-AETHERMOORE hyperbolic geometry AI safety and governance framework.
Overview
This dataset contains 10,978 training pairs spanning the full SCBE-AETHERMOORE system: 14-layer architecture knowledge, Six Sacred… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-aethermoore-training-data.scbe-system-hygiene-training-data
Status: experimental. Experiment-specific slice. Primary public dataset: scbe-aethermoore-training-data.
SCBE System Hygiene Training
Metadata-only SCBE cleanup training records generated at 2026-05-12T06:55:13Z.
This dataset teaches local-first cleanup decisions: keep harness-wired models,
review ambiguous model/cache state, and turn deletion candidates into scrubbed
training examples before pruning local storage.
It does not include raw cache files, model weights, local logs… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-system-hygiene-training-data.scbe-codeflow-bijective-v1
Status: experimental. Experiment-specific slice. Primary public dataset: scbe-aethermoore-training-data.
SCBE Codeflow Bijective v1
Supervised fine-tuning corpus teaching bijective multi-tongue / multi-language
code editing. Each algorithm is decomposed into N semantic slots. Every slot
is filled in all 6 Sacred Tongues. An edit at slot k in any tongue maps
deterministically to the parallel slot k in every other tongue. Syntactic line
counts may differ per tongue; semantic… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-codeflow-bijective-v1.scbe-life-science-research-training-demo
Status: experimental. Experiment-specific slice. Primary public dataset: scbe-aethermoore-training-data.
SCBE Research Training Package
This package was generated from live pubmed pulls for the query protein structure prediction and is meant for
lightweight Hugging Face dataset and SFT experiments.
Files
papers.jsonl: normalized raw research records
sft_train.jsonl: train split for instruction-style tasks
sft_validation.jsonl: validation split… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-life-science-research-training-demo.scbe-tongue-drill-sft-v1
Status: experimental. Experiment-specific slice. Primary public dataset: scbe-aethermoore-training-data.
SCBE Tongue Drill SFT v1
Supervised fine-tuning drill dataset for the SCBE Sacred Tongues table-lock system.
Each row is a 3-turn chat (system / user / assistant) teaching the model to emit
canonical packets verbatim for a given (map, tongue, value) triple.
Splits
Split
Rows
all
2630
train
2373
holdout
257
Holdout is row_index % 10 == 0… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-tongue-drill-sft-v1.
