research-results
llama2_7b_chat-boolq-results
Dataset Card for "llama2_7b_chat-boolq-results"
More Information needed
lca-results
Long Code Arena (raw results)
These are the raw results from the Long Code Arena benchmark suite, as well as the corresponding model predictions.
Please use the subset dropdown menu to select the necessary data relating to our six benchmarks:
🤗 Library-based code generation
🤗 CI builds repair
🤗 Project-level code completion
🤗 Commit message generation🤗 Bug localization
🤗 Module summarization
llama2_7b_chat-piqa-resultsdementor-complete-experiment-results
Dementor complete experiment results
Audited outputs for the configuration-defined Dementor completion campaign.
Audited scope
Behavioral imitation adapters: 1,104 total (528 SFT, 528 DPO, 48 self-SFT controls).
Behavioral-fidelity evaluation: 1,104 adapters on 200 held-out prompts, with embedding and
primary LLM-judge scores, plus 48 target-reference response sets.
Activation steering: 29 models, seven benchmarks, and two operators (original and fpall),
totaling… See the full description on the dataset page: https://huggingface.co/datasets/dementor-research/dementor-complete-experiment-results.phi-winogrande_inverted_option-results
Dataset Card for "phi-winogrande_inverted_option-results"
More Information needed
Research_results
