full-context
CANBERT_-_phi3-full-context-ggufqwen3-4b-instruct-2507-eagle3-sharegpt-full-context-epoch0-step15000qwen3-4b-instruct-2507-eagle3-sharegpt-full-context-epoch7-step185000qwen3-4b-instruct-2507-eagle3-sharegpt-full-context-epoch2-step60000qwen3-4b-instruct-2507-eagle3-sharegpt-full-context-epoch6-step155000qwen3-4b-instruct-2507-eagle3-sharegpt-full-context-epoch4-step110000qwen3-4b-instruct-2507-eagle3-sharegpt-full-context-epoch0-step10000qwen3-4b-instruct-2507-eagle3-sharegpt-full-context-epoch1-step30000
vla0-context-trace-full-demo-datasetqwen3-thinking-full-contexts-8kfineproofs-prm-context-v2-full-cot-xprob
FineProofs PRM Context v2: Full Cot Xprob
This arm uses full_cot context from cross_problem rollouts with packing policy middle_truncated_reasoning_equal_share.
This is one of nine row-matched context variants built from the verified FineProofs rollout collection. Partial-prefix and complete-response targets both use canonical normalized rubric credit derived from clamped points divided by max points. The correct column is only a legacy boolean projection at reward >= 0.5;… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/fineproofs-prm-context-v2-full-cot-xprob.prm-mc-value-context-full-cot
MC-value PRM dataset — context mode: full_cot
Data for training an in-context / policy-conditioned Monte-Carlo value PRM. Each row's query is a
partial reasoning prefix; the target reward = V = P(correct | prefix), the Monte-Carlo value estimated
from branched Qwen3.5-4B rollouts on Polaris math problems. The user prompt additionally carries a
"# Other attempts by the same model at this problem" block — the ablation variable.
Context for this variant: Up to K=4 OTHER attempts'… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/prm-mc-value-context-full-cot.fineproofs-prm-context-v2-full-cot
FineProofs PRM Context v2: Full Cot
This arm uses full_cot context from same_problem rollouts with packing policy middle_truncated_reasoning_equal_share.
This is one of nine row-matched context variants built from the verified FineProofs rollout collection. Partial-prefix and complete-response targets both use canonical normalized rubric credit derived from clamped points divided by max points. The correct column is only a legacy boolean projection at reward >= 0.5; training… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/fineproofs-prm-context-v2-full-cot.superskillret-index-fullcontext
superskillret prebuilt index — full-context
Prebuilt embedding index for the superskillret Claude Code plugin.
Unlike the default index (which embeds only name + description), this build
encodes the full skill body (name + description + body) up to
max_seq_length=32768 tokens. Larger index, much higher recall on
skills whose name/description don't capture every keyword in the body.
Version: 1
Corpus: ThakiCloud/SKILLRET (train+test)
Encoder: ThakiCloud/SkillRet-Embedding-0.6B… See the full description on the dataset page: https://huggingface.co/datasets/youngryankim/superskillret-index-fullcontext.
