MarioBarbeque/CyberSolve-LinAlg-1.2-correctness-benchmark
112
dataset constructed for benchmarking the partial correctness of our finetuned CyberSolve-LingAlg-1.2 model
initial commit
dataset constructed for benchmarking the partial correctness of our finetuned CyberSolve-LingAlg-1.2 model
initial commit