CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01malaiwah /spark2-5-tiny-fidelity-root-v1 spark2-5 random CPU fixture root A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/spark2-5-tiny-random-bf16. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/spark2-5-tiny-fidelity-root-v1.tabularn<1K0 likes70 downloads16d agoHugging Face02open-llm-leaderboard /arcee-ai__Llama-Spark-detailsgated Dataset Card for Evaluation run of arcee-ai/Llama-Spark Dataset automatically created during the evaluation run of model arcee-ai/Llama-Spark The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/arcee-ai__Llama-Spark-details.tabular10K<n<100K0 likes47 downloads2y agoHugging Face03aaaded /spark-capacity-boundary-study Capacity-boundary optimization with Spark-X2.5-1.7B This project contains an original evaluation for HER Hack-Astron #6. The report is in DISCUSSION.md. Creating or publishing these artifacts is not an award or payment. The experiment checks whether increasing a six-item 0/1 knapsack's capacity by one causes the model to find the new optimum. Four seeded item families produce eight mathematical instances, each in English and Chinese. Every prompt runs once with thinking off and… See the full description on the dataset page: https://huggingface.co/datasets/aaaded/spark-capacity-boundary-study.tabularn<1K0 likes47 downloads12d agoHugging Face04sparksofagi /mhpp MHPP (questions only) Evaluation-only questions for code/problem-solving. Do not train on test. Data fields id (int), question (string), prompt (string), function_name (string), parameters (list[str]), difficulty_types (int) tabularn<1K0 likes36 downloads1y agoHugging Face05open-llm-leaderboard /Jacoby746__Proto-Harpy-Spark-v0.1-7B-detailsgated Dataset Card for Evaluation run of Jacoby746/Proto-Harpy-Spark-v0.1-7B Dataset automatically created during the evaluation run of model Jacoby746/Proto-Harpy-Spark-v0.1-7B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jacoby746__Proto-Harpy-Spark-v0.1-7B-details.tabular10K<n<100K0 likes34 downloads2y agoHugging Face06sparkthai /sparkthai-model-arena SparkThai Model Arena Question bank, versioned snapshots, run history, and full per-trial results from the SparkThai Model Arena — SparkThai's own language-model eval suite, run on our own NVIDIA DGX Spark hardware. 🌐 sparkthai.com · 🏆 Model Arena · 🐙 github.com/sparkthai What this is The Model Arena grades Qwen model variants we host ourselves on DGX Spark across four kinds of checks: code_exact — code tasks graded by running the model's code against an… See the full description on the dataset page: https://huggingface.co/datasets/sparkthai/sparkthai-model-arena.tabularn<1K0 likes22 downloads1mo agoHugging Face07open-llm-leaderboard /arcee-ai__Arcee-Spark-detailsgated Dataset Card for Evaluation run of arcee-ai/Arcee-Spark Dataset automatically created during the evaluation run of model arcee-ai/Arcee-Spark The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/arcee-ai__Arcee-Spark-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.