CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Stage-jh-monitor /appworld-qwen35-4b-9b-s_signal_6-epoch4-iter1 appworld-qwen35-4b-9b-s_signal_6-epoch4-iter1 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3953125 Action score: 0.446875 Valid samples: 320/320 tabularn<1K0 likes4.1k downloads15d agoHugging Face02Stage-jh-monitor /appworld-qwen35-4b-9b-s_signal_5-epoch4-iter1-reeval1 appworld-qwen35-4b-9b-s_signal_5-epoch4-iter1-reeval1 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.41328125 Action score: 0.4359375 Valid samples: 320/320 tabularn<1K0 likes3.7k downloads13d agoHugging Face03Stage-jh-monitor /appworld-qwen35-4b-9b-s_signal_5-epoch4-iter1 appworld-qwen35-4b-9b-s_signal_5-epoch4-iter1 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.41953125 Action score: 0.4515625 Valid samples: 320/320 tabularn<1K0 likes3.7k downloads13d agoHugging Face04Dogacel /nemotron-post-training-v2-qwen-3.5-9b-regen Dataset Card for Nemotron Post Training v2 Qwen 3.5 9B Regen Regenerated responses from nvidia/Nemotron-Post-Training-Dataset-v2 dataset using Qwen3.5 9B model. Parameter Value Max Tokens 4096 Temperature 1.0 Top-k 20 Top-p 0.95 Repetition Penalty 1.5 Dataset consists only the english samples from the Nemotron Post Training Dataset. 85% of the chat prompts have reasoning enabled, every other category has reasoning disabled. Category Value math… See the full description on the dataset page: https://huggingface.co/datasets/Dogacel/nemotron-post-training-v2-qwen-3.5-9b-regen.texttext-generation100K<n<1M0 likes1.8k downloads5mo agoHugging Face05Ouroboros-Research /llama-9b-bulk-npztabularn<1K0 likes541 downloads15d agoHugging Face06CK0607 /qwen3.5-9b-blogprovider-traces Blog-Provider-ID — model inference traces (val + val_ood) Per-model generation traces for the 3-way AI-provider classification task (CLAUDE / CHATGPT / GEMINI), produced by the models in the CK0607 blog-provider collection. Each model folder holds val.jsonl, val_ood.jsonl (one record per blog: prompt gold, prediction, full <reason_why>/<answer> completion, truncation flag) and a summary.json. All inference used the plain SYSTEM_PROMPT_3WAY (thinking OFF), so numbers are directly… See the full description on the dataset page: https://huggingface.co/datasets/CK0607/qwen3.5-9b-blogprovider-traces.tabulartext-classification1K<n<10K1 likes84 downloads3mo agoHugging Face07open-llm-leaderboard /01-ai__Yi-1.5-9B-Chat-detailsgated Dataset Card for Evaluation run of 01-ai/Yi-1.5-9B-Chat Dataset automatically created during the evaluation run of model 01-ai/Yi-1.5-9B-Chat The dataset is composed of 78 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-1.5-9B-Chat-details.tabular10K<n<100K0 likes76 downloads2y agoHugging Face08axiomofmind /Angry-Claudius-9B-Dataset Angry Claudius 9B Dataset The training and evaluation data used to develop Angry Claudius 9B, a joke model trained to answer user requests with short profane dismissals instead of completing the requested task. Content warning This dataset contains frequent explicit profanity. It is intended for behavioral fine-tuning and evaluation research and is unsuitable for applications that require polite, helpful, or family-friendly responses. Data… See the full description on the dataset page: https://huggingface.co/datasets/axiomofmind/Angry-Claudius-9B-Dataset.texttext-generation1K<n<10K0 likes76 downloads13d agoHugging Face09open-llm-leaderboard /01-ai__Yi-1.5-9B-detailsgated Dataset Card for Evaluation run of 01-ai/Yi-1.5-9B Dataset automatically created during the evaluation run of model 01-ai/Yi-1.5-9B The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-1.5-9B-details.tabular10K<n<100K0 likes68 downloads2y agoHugging Face10open-llm-leaderboard /zelk12__MT2-Gen7-gemma-2-9B-detailsgated Dataset Card for Evaluation run of zelk12/MT2-Gen7-gemma-2-9B Dataset automatically created during the evaluation run of model zelk12/MT2-Gen7-gemma-2-9B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/zelk12__MT2-Gen7-gemma-2-9B-details.tabular10K<n<100K0 likes68 downloads2y agoHugging Face11visual-memory /ConvAI2-Qwen-enhanced-Qwen3.5-9B Visual Memory Results: convai2-qwen-enhanced This dataset contains the scored output of a visual-memory perplexity experiment. Experiment metadata { "experiment": { "model_name": "Qwen/Qwen3.5-9B", "hf_results_repo": "visual-memory/ConvAI2-Qwen-enhanced-Qwen3.5-9B", "results_jsonl": "results/ConvAI2-Qwen-enhanced-Qwen3.5-9B.jsonl", "hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy", "hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-enhanced-Qwen3.5-9B.tabular1K<n<10K0 likes54 downloads7d agoHugging Face12malaiwah /qfs-capture-9bafb145820e HF workflow 9bafb145820e0eba1117bbfcefff8abc A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/qwen3-5-tiny-random-gptq-v1-g32-rtn-format. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-capture-9bafb145820e.tabularn<1K0 likes53 downloads16d agoHugging Face13visual-memory /Synthetic-Persona-Chat-Qwen-enhanced-Qwen3.5-9B Visual Memory Results: synthetic-persona-chat-qwen-enhanced This dataset contains the scored output of a visual-memory perplexity experiment. Experiment metadata { "experiment": { "model_name": "Qwen/Qwen3.5-9B", "hf_results_repo": "visual-memory/Synthetic-Persona-Chat-Qwen-enhanced-Qwen3.5-9B", "results_jsonl": "results/Synthetic-Persona-Chat-Qwen-enhanced-Qwen3.5-9B.jsonl", "hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-Qwen-enhanced-Qwen3.5-9B.tabular1K<n<10K0 likes52 downloads7d agoHugging Face14visual-memory /Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-9B Visual Memory Results: synthetic-persona-chat-flux-enhanced This dataset contains the scored output of a visual-memory perplexity experiment. Experiment metadata { "experiment": { "model_name": "Qwen/Qwen3.5-9B", "hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-9B", "results_jsonl": "results/Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-9B.jsonl", "hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-9B.tabular1K<n<10K0 likes50 downloads7d agoHugging Face15visual-memory /Synthetic-Persona-Chat-ERNIE-original-Qwen3.5-9B Visual Memory Results: synthetic-persona-chat-ernie-original This dataset contains the scored output of a visual-memory perplexity experiment. Experiment metadata { "experiment": { "model_name": "Qwen/Qwen3.5-9B", "hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-original-Qwen3.5-9B", "results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-original-Qwen3.5-9B.jsonl", "hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-original-Qwen3.5-9B.tabular1K<n<10K0 likes50 downloads7d agoHugging Face16open-llm-leaderboard /nhyha__N3N_gemma-2-9b-it_20241029_1532-detailsgated Dataset Card for Evaluation run of nhyha/N3N_gemma-2-9b-it_20241029_1532 Dataset automatically created during the evaluation run of model nhyha/N3N_gemma-2-9b-it_20241029_1532 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nhyha__N3N_gemma-2-9b-it_20241029_1532-details.tabular10K<n<100K0 likes49 downloads2y agoHugging Face17visual-memory /Synthetic-Persona-Chat-ERNIE-enhanced-Qwen3.5-9B Visual Memory Results: synthetic-persona-chat-ernie-enhanced This dataset contains the scored output of a visual-memory perplexity experiment. Experiment metadata { "experiment": { "model_name": "Qwen/Qwen3.5-9B", "hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-Qwen3.5-9B", "results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-enhanced-Qwen3.5-9B.jsonl", "hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-Qwen3.5-9B.tabular1K<n<10K0 likes49 downloads7d agoHugging Face18visual-memory /Synthetic-Persona-Chat-Qwen-original-Qwen3.5-9B Visual Memory Results: synthetic-persona-chat-qwen-original This dataset contains the scored output of a visual-memory perplexity experiment. Experiment metadata { "experiment": { "model_name": "Qwen/Qwen3.5-9B", "hf_results_repo": "visual-memory/Synthetic-Persona-Chat-Qwen-original-Qwen3.5-9B", "results_jsonl": "results/Synthetic-Persona-Chat-Qwen-original-Qwen3.5-9B.jsonl", "hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-Qwen-original-Qwen3.5-9B.tabular1K<n<10K0 likes48 downloads7d agoHugging Face19open-llm-leaderboard /Quazim0t0__Mouse-9B-detailsgated Dataset Card for Evaluation run of Quazim0t0/Mouse-9B Dataset automatically created during the evaluation run of model Quazim0t0/Mouse-9B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Quazim0t0__Mouse-9B-details.tabular10K<n<100K0 likes47 downloads2y agoHugging Face20visual-memory /ConvAI2-ERNIE-original-Qwen3.5-9B Visual Memory Results: convai2-ernie-original This dataset contains the scored output of a visual-memory perplexity experiment. Experiment metadata { "experiment": { "model_name": "Qwen/Qwen3.5-9B", "hf_results_repo": "visual-memory/ConvAI2-ERNIE-original-Qwen3.5-9B", "results_jsonl": "results/ConvAI2-ERNIE-original-Qwen3.5-9B.jsonl", "hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy", "hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-original-Qwen3.5-9B.tabular1K<n<10K0 likes47 downloads7d agoHugging Face21visual-memory /Synthetic-Persona-Chat-FLUX-original-Qwen3.5-9B Visual Memory Results: synthetic-persona-chat-flux-original This dataset contains the scored output of a visual-memory perplexity experiment. Experiment metadata { "experiment": { "model_name": "Qwen/Qwen3.5-9B", "hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-original-Qwen3.5-9B", "results_jsonl": "results/Synthetic-Persona-Chat-FLUX-original-Qwen3.5-9B.jsonl", "hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-original-Qwen3.5-9B.tabular1K<n<10K0 likes47 downloads7d agoHugging Face22open-llm-leaderboard /lemon07r__Gemma-2-Ataraxy-v4-Advanced-9B-detailsgated Dataset Card for Evaluation run of lemon07r/Gemma-2-Ataraxy-v4-Advanced-9B Dataset automatically created during the evaluation run of model lemon07r/Gemma-2-Ataraxy-v4-Advanced-9B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lemon07r__Gemma-2-Ataraxy-v4-Advanced-9B-details.tabular10K<n<100K0 likes46 downloads2y agoHugging Face23violetxi /ch-pilot-rollouts-qwen3.5-9b C&H Pilot Rollouts — Qwen/Qwen3.5-9B 20 agentic exploration rollouts over the Calderwood & Harkness (C&H) synthetic law-firm corpus (the open-sourced world from harvey-labs tasks/firm-knowledge/, MIT), generated by Qwen/Qwen3.5-9B served with vLLM. Part of an actor-selection pilot for a world-internalization research project: the goal is to mine agent trajectories into verified fact stores and rewritten likelihood-training targets. Companion dataset (same seeds/tasks, different… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/ch-pilot-rollouts-qwen3.5-9b.tabulartext-generationn<1K0 likes46 downloads1mo agoHugging Face24visual-memory /PersonaChat-Qwen-original-Qwen3.5-9B Visual Memory Results: personachat-qwen-original This dataset contains the scored output of a visual-memory perplexity experiment. Experiment metadata { "experiment": { "model_name": "Qwen/Qwen3.5-9B", "hf_results_repo": "visual-memory/PersonaChat-Qwen-original-Qwen3.5-9B", "results_jsonl": "results/PersonaChat-Qwen-original-Qwen3.5-9B.jsonl", "hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy", "hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-Qwen-original-Qwen3.5-9B.tabular1K<n<10K0 likes46 downloads8d agoHugging Face25visual-memory /ConvAI2-FLUX-original-Qwen3.5-9B Visual Memory Results: convai2-flux-original This dataset contains the scored output of a visual-memory perplexity experiment. Experiment metadata { "experiment": { "model_name": "Qwen/Qwen3.5-9B", "hf_results_repo": "visual-memory/ConvAI2-FLUX-original-Qwen3.5-9B", "results_jsonl": "results/ConvAI2-FLUX-original-Qwen3.5-9B.jsonl", "hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy", "hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-original-Qwen3.5-9B.tabular1K<n<10K0 likes46 downloads7d agoHugging Face26visual-memory /ConvAI2-FLUX-enhanced-Qwen3.5-9B Visual Memory Results: convai2-flux-enhanced This dataset contains the scored output of a visual-memory perplexity experiment. Experiment metadata { "experiment": { "model_name": "Qwen/Qwen3.5-9B", "hf_results_repo": "visual-memory/ConvAI2-FLUX-enhanced-Qwen3.5-9B", "results_jsonl": "results/ConvAI2-FLUX-enhanced-Qwen3.5-9B.jsonl", "hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy", "hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-enhanced-Qwen3.5-9B.tabular1K<n<10K0 likes46 downloads7d agoHugging Face27open-llm-leaderboard /zelk12__MTM-Merge-gemma-2-9B-detailsgated Dataset Card for Evaluation run of zelk12/MTM-Merge-gemma-2-9B Dataset automatically created during the evaluation run of model zelk12/MTM-Merge-gemma-2-9B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/zelk12__MTM-Merge-gemma-2-9B-details.tabular10K<n<100K0 likes45 downloads2y agoHugging Face28KermitCO /qwen3.5-9B-tau2bench-retail-baseline-traces Qwen3.5-9B tau2-bench retail baseline traces (n=3 × 114) Three independent evaluation trials of Qwen3.5-9B (no fine-tune, no memory) on the full 114-task tau2-bench retail set. Each trial_N.jsonl is one trial; one JSON object per line, one object per task. Per-trial pass^1 (canonical reward) trial_1: 72.8% trial_2: 69.3% trial_3: 73.7% Pooled metrics (n=3 across 114 tasks) pass^1 = 71.9% pass^2 = 58.2% pass^3 = 49.1% Trace fields (per object)… See the full description on the dataset page: https://huggingface.co/datasets/KermitCO/qwen3.5-9B-tau2bench-retail-baseline-traces.tabularn<1K0 likes45 downloads4mo agoHugging Face29dongboklee /MMLU-Pro_gemma-2-9b-it_test MMLU-Pro_gemma-2-9b-it_test text1K<n<10K0 likes45 downloads3mo agoHugging Face30visual-memory /PersonaChat-ERNIE-enhanced-Qwen3.5-9B Visual Memory Results: personachat-ernie-enhanced This dataset contains the scored output of a visual-memory perplexity experiment. Experiment metadata { "experiment": { "model_name": "Qwen/Qwen3.5-9B", "hf_results_repo": "visual-memory/PersonaChat-ERNIE-enhanced-Qwen3.5-9B", "results_jsonl": "results/PersonaChat-ERNIE-enhanced-Qwen3.5-9B.jsonl", "hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy", "hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-ERNIE-enhanced-Qwen3.5-9B.tabular1K<n<10K0 likes45 downloads8d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.