sparkx2_5
spark-x25-congruence-study
A one-unit change that removes every solution
This repository contains an original 32-problem linear-congruence corpus, 64 real Spark-X2.5-1.7B BF16 responses, and reproducible evaluation code. Each of 16 solvable problems is paired with an unsolvable problem whose right-hand side differs by one. Both thinking settings were run once under a 2,048-token cap.
Strict full-answer scores were 17/32 without thinking and 27/32 with thinking. These are results on a small constructed… See the full description on the dataset page: https://huggingface.co/datasets/zhengjiyue/spark-x25-congruence-study.spark-x25-1.7b-gsm8k-raw-archive
Spark-X2.5-1.7B GSM8K Evaluation — Raw Data Archive
Reproducibility archive for https://huggingface.co/XHToken/Spark-X2.5-1.7B/discussions/38
Contents
File
What it is
eval2.py
Original 0-shot harness (200-problem seed-42 sample, llama.cpp server, strict #### <number> scoring)
greedy_retry.py
Token-budget retry harness for the 59 misses (identical prompts/sampling, max_tokens=3000)
results.jsonl
Original 0-shot run: 200 problems + 95 connection-retry… See the full description on the dataset page: https://huggingface.co/datasets/Hemant95/spark-x25-1.7b-gsm8k-raw-archive.spark-x25-integrality-evaluation
Reproducible evidence
Download the complete evidence ZIP. It contains all 72 complete outputs and token IDs, annotations, exact gold, code, protocols, and reproduction instructions.
ZIP SHA256: ef27046e757893d5c10107efabaefda3d93c08d2007292e4d498c9c0755d12ea. Size: 399,106 bytes. 101 files, including a manifest covering the other 100 files.
[HER Hack-Astron #6] Whole runs or fractional runs? Spark-X2.5 production planning under a fixed output budget
This is an… See the full description on the dataset page: https://huggingface.co/datasets/retrainmap/spark-x25-integrality-evaluation.spark-x25-bilingual-math-evaluation
Spark-X2.5 bilingual math and output-budget study
Prepared for HER Hack-Astron #6.
This is an AI-operated local experiment for the human account owner (Hugging Face jojoqiao940, GitHub shiyingqiao940-alt). Codex designed the probes, wrote and ran code, inspected solutions, and drafted the report. No independent human adjudication is claimed. The account owner reviewed the submission brief and approved publication. No independent human adjudication, award or payment is claimed.… See the full description on the dataset page: https://huggingface.co/datasets/jojoqiao940/spark-x25-bilingual-math-evaluation.
