datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scbe-codeflow-bijective-v1
Status: experimental. Experiment-specific slice. Primary public dataset: scbe-aethermoore-training-data.
SCBE Codeflow Bijective v1
Supervised fine-tuning corpus teaching bijective multi-tongue / multi-language
code editing. Each algorithm is decomposed into N semantic slots. Every slot
is filled in all 6 Sacred Tongues. An edit at slot k in any tongue maps
deterministically to the parallel slot k in every other tongue. Syntactic line
counts may differ per tongue; semantic… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-codeflow-bijective-v1.scbe-life-science-research-training-demo
Status: experimental. Experiment-specific slice. Primary public dataset: scbe-aethermoore-training-data.
SCBE Research Training Package
This package was generated from live pubmed pulls for the query protein structure prediction and is meant for
lightweight Hugging Face dataset and SFT experiments.
Files
papers.jsonl: normalized raw research records
sft_train.jsonl: train split for instruction-style tasks
sft_validation.jsonl: validation split… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-life-science-research-training-demo.scbe-tongue-drill-sft-v1
Status: experimental. Experiment-specific slice. Primary public dataset: scbe-aethermoore-training-data.
SCBE Tongue Drill SFT v1
Supervised fine-tuning drill dataset for the SCBE Sacred Tongues table-lock system.
Each row is a 3-turn chat (system / user / assistant) teaching the model to emit
canonical packets verbatim for a given (map, tongue, value) triple.
Splits
Split
Rows
all
2630
train
2373
holdout
257
Holdout is row_index % 10 == 0… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-tongue-drill-sft-v1.
