jang
Datasets
All datasets matching “jang”SCBench-preprocessedThis is the preprocessed version of Microsoft SCBench, used by KVzip:
Each data example has a format of {context: str, question: List[str], answers: List[str]}
Each dataset contains only examples whose context token length (measured with the LLaMA3 tokenizer) is less than 125K, fitting within the context limit of LLaMA3 models.
We also provide shortened versions of SCBench, excluding tasks {choice_eng, qa_eng, and vt}, which are difficult to shorten.
The "tiny" tag (e.g., scbench_kv_tiny)… See the full description on the dataset page: https://huggingface.co/datasets/Jang-Hyun/SCBench-preprocessed.agentic-drug-discovery-system
Agentic Drug Discovery System
This card describes the public 0.3.0.dev3 Agentic Drug Discovery System mirror.
Scope. The proposed eight-stage, long-horizon agentic drug discovery system remains a research scaffold rather than a completed public platform. Seven of eight planned atlases have no standalone public data, and the demonstrated continuous multi-stage program currently covers one disease/target slice traversed retrospectively.
It contains the executable control plane… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/agentic-drug-discovery-system.narrow-model-safety-eval
Narrow Model Safety Evaluation — Protein Dual-Use Risk Dataset
Summary: Annotations, results, and evaluation data for a proof-of-concept framework assessing dual-use risk in narrow scientific AI models. Two lines of work: (1) structure-level metrics — FSPE, FSI, and Physical Realizability Tier — on eight published protein toxins and mechanism-matched benign controls (ESM-2, ProteinMPNN); (2) mechanism generalization — a leave-one-mechanism-out panel measuring what an… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/narrow-model-safety-eval.teacher-embedding-corpusnepali-corpusThis is a clone of https://huggingface.co/datasets/Boredoom17/Nepali-Corpus but the data is sharded for efficiency.
grounding-atlas
grounding-atlas: verifiable-signal pairs
Matched (representation, verifiable-property) pairs for measuring whether a
language model grounds the content of a scientific representation (a SMILES
string, a protein/DNA/RNA sequence, an expression vector, a spectrum, an image)
or merely its name. Each property is either an experimentally measured endpoint
or a closed-form function of the representation, so the representation is the
ground truth and grounding becomes directly… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/grounding-atlas.
