bakeoff
Datasets
All datasets matching “bakeoff”exp10-deploy-bakeoff-data
Exp10 — Constrained-Choice Steering for Eval-Awareness
Refining eval-awareness probes via gradient descent against constrained behavioral
losses (MCQ today; fill-in-the-blank and FORTRESS-vs-WildChat distinguishability
planned). Tests whether contrastive probes capture self-locating belief or just
user-asserted context.
Two model organisms:
M_sdf — nvidia/Llama-3_3-Nemotron-Super-49B-v1 + Wood Labs SDF LoRA. Type-hint usage is the primary behavioral signal.
OLMO 3.1 32B Think —… See the full description on the dataset page: https://huggingface.co/datasets/rlundqvist/exp10-deploy-bakeoff-data.dmb-bakeoff-hardLLM_reasoning_bakeofflong-context-baseline-bakeoff
Long-Context Data-Selection Bake-off — Shared Candidate Pool
The shared 16K candidate pool for comparing long-context data-selection methods on equal
footing. Every method (AttentionSpan, LongAttn, LongProc, ProLong, perplexity, ...) scores the
same 14,300 documents, picks its own top-800 under the same split, then trains
Llama-2-7B + 16K LoRA and evaluates on HELMET.
Files
File
Description
candidate_pool_16k_scored.parquet
The shared pool — 14,300 docs… See the full description on the dataset page: https://huggingface.co/datasets/KevinDavidHayes/long-context-baseline-bakeoff.hucker-bakeoff2-deepseek2-ordinary
Document OCR using DeepSeek-OCR-2
This dataset contains markdown-formatted OCR results from images in bokane/hucker-ocr-bakeoff using DeepSeek-OCR-2.
Processing Details
Source Dataset: bokane/hucker-ocr-bakeoff
Model: deepseek-ai/DeepSeek-OCR-2
Number of Samples: 12
Processing Time: 2.4 min
Processing Date: 2026-08-13 22:40 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split: train
Batch Size: 8
Max Model Length: 8,192… See the full description on the dataset page: https://huggingface.co/datasets/bokane/hucker-bakeoff2-deepseek2-ordinary.dmb-bakeoff-pages
