datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
appworld-qwen35-4b-9b-s_signal_6-epoch4-iter1
appworld-qwen35-4b-9b-s_signal_6-epoch4-iter1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3953125
Action score: 0.446875
Valid samples: 320/320
appworld-qwen35-4b-9b-s_signal_5-epoch4-iter1-reeval1
appworld-qwen35-4b-9b-s_signal_5-epoch4-iter1-reeval1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.41328125
Action score: 0.4359375
Valid samples: 320/320
appworld-qwen35-4b-9b-s_signal_5-epoch4-iter1
appworld-qwen35-4b-9b-s_signal_5-epoch4-iter1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.41953125
Action score: 0.4515625
Valid samples: 320/320
llama-9b-bulk-npzqwen3.5-9b-blogprovider-traces
Blog-Provider-ID — model inference traces (val + val_ood)
Per-model generation traces for the 3-way AI-provider classification task (CLAUDE / CHATGPT / GEMINI),
produced by the models in the CK0607 blog-provider collection.
Each model folder holds val.jsonl, val_ood.jsonl (one record per blog: prompt gold, prediction,
full <reason_why>/<answer> completion, truncation flag) and a summary.json.
All inference used the plain SYSTEM_PROMPT_3WAY (thinking OFF), so numbers are directly… See the full description on the dataset page: https://huggingface.co/datasets/CK0607/qwen3.5-9b-blogprovider-traces.01-ai__Yi-1.5-9B-Chat-details
Dataset Card for Evaluation run of 01-ai/Yi-1.5-9B-Chat
Dataset automatically created during the evaluation run of model 01-ai/Yi-1.5-9B-Chat
The dataset is composed of 78 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-1.5-9B-Chat-details.01-ai__Yi-1.5-9B-details
Dataset Card for Evaluation run of 01-ai/Yi-1.5-9B
Dataset automatically created during the evaluation run of model 01-ai/Yi-1.5-9B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-1.5-9B-details.zelk12__MT2-Gen7-gemma-2-9B-details
Dataset Card for Evaluation run of zelk12/MT2-Gen7-gemma-2-9B
Dataset automatically created during the evaluation run of model zelk12/MT2-Gen7-gemma-2-9B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/zelk12__MT2-Gen7-gemma-2-9B-details.ConvAI2-Qwen-enhanced-Qwen3.5-9B
Visual Memory Results: convai2-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-9B",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-enhanced-Qwen3.5-9B",
"results_jsonl": "results/ConvAI2-Qwen-enhanced-Qwen3.5-9B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-enhanced-Qwen3.5-9B.qfs-capture-9bafb145820e
HF workflow 9bafb145820e0eba1117bbfcefff8abc
A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/qwen3-5-tiny-random-gptq-v1-g32-rtn-format.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-capture-9bafb145820e.Synthetic-Persona-Chat-Qwen-enhanced-Qwen3.5-9B
Visual Memory Results: synthetic-persona-chat-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-9B",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-Qwen-enhanced-Qwen3.5-9B",
"results_jsonl": "results/Synthetic-Persona-Chat-Qwen-enhanced-Qwen3.5-9B.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-Qwen-enhanced-Qwen3.5-9B.Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-9B
Visual Memory Results: synthetic-persona-chat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-9B",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-9B",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-9B.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-9B.Synthetic-Persona-Chat-ERNIE-original-Qwen3.5-9B
Visual Memory Results: synthetic-persona-chat-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-9B",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-original-Qwen3.5-9B",
"results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-original-Qwen3.5-9B.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-original-Qwen3.5-9B.nhyha__N3N_gemma-2-9b-it_20241029_1532-details
Dataset Card for Evaluation run of nhyha/N3N_gemma-2-9b-it_20241029_1532
Dataset automatically created during the evaluation run of model nhyha/N3N_gemma-2-9b-it_20241029_1532
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nhyha__N3N_gemma-2-9b-it_20241029_1532-details.Synthetic-Persona-Chat-ERNIE-enhanced-Qwen3.5-9B
Visual Memory Results: synthetic-persona-chat-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-9B",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-Qwen3.5-9B",
"results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-enhanced-Qwen3.5-9B.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-Qwen3.5-9B.Synthetic-Persona-Chat-Qwen-original-Qwen3.5-9B
Visual Memory Results: synthetic-persona-chat-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-9B",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-Qwen-original-Qwen3.5-9B",
"results_jsonl": "results/Synthetic-Persona-Chat-Qwen-original-Qwen3.5-9B.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-Qwen-original-Qwen3.5-9B.Quazim0t0__Mouse-9B-details
Dataset Card for Evaluation run of Quazim0t0/Mouse-9B
Dataset automatically created during the evaluation run of model Quazim0t0/Mouse-9B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Quazim0t0__Mouse-9B-details.ConvAI2-ERNIE-original-Qwen3.5-9B
Visual Memory Results: convai2-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-9B",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-original-Qwen3.5-9B",
"results_jsonl": "results/ConvAI2-ERNIE-original-Qwen3.5-9B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-original-Qwen3.5-9B.Synthetic-Persona-Chat-FLUX-original-Qwen3.5-9B
Visual Memory Results: synthetic-persona-chat-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-9B",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-original-Qwen3.5-9B",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-original-Qwen3.5-9B.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-original-Qwen3.5-9B.lemon07r__Gemma-2-Ataraxy-v4-Advanced-9B-details
Dataset Card for Evaluation run of lemon07r/Gemma-2-Ataraxy-v4-Advanced-9B
Dataset automatically created during the evaluation run of model lemon07r/Gemma-2-Ataraxy-v4-Advanced-9B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lemon07r__Gemma-2-Ataraxy-v4-Advanced-9B-details.ch-pilot-rollouts-qwen3.5-9b
C&H Pilot Rollouts — Qwen/Qwen3.5-9B
20 agentic exploration rollouts over the Calderwood & Harkness (C&H) synthetic law-firm
corpus (the open-sourced world from harvey-labs
tasks/firm-knowledge/, MIT), generated by Qwen/Qwen3.5-9B served with vLLM.
Part of an actor-selection pilot for a world-internalization research project: the goal is to
mine agent trajectories into verified fact stores and rewritten likelihood-training targets.
Companion dataset (same seeds/tasks, different… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/ch-pilot-rollouts-qwen3.5-9b.PersonaChat-Qwen-original-Qwen3.5-9B
Visual Memory Results: personachat-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-9B",
"hf_results_repo": "visual-memory/PersonaChat-Qwen-original-Qwen3.5-9B",
"results_jsonl": "results/PersonaChat-Qwen-original-Qwen3.5-9B.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-Qwen-original-Qwen3.5-9B.ConvAI2-FLUX-original-Qwen3.5-9B
Visual Memory Results: convai2-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-9B",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-original-Qwen3.5-9B",
"results_jsonl": "results/ConvAI2-FLUX-original-Qwen3.5-9B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-original-Qwen3.5-9B.ConvAI2-FLUX-enhanced-Qwen3.5-9B
Visual Memory Results: convai2-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-9B",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-enhanced-Qwen3.5-9B",
"results_jsonl": "results/ConvAI2-FLUX-enhanced-Qwen3.5-9B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-enhanced-Qwen3.5-9B.zelk12__MTM-Merge-gemma-2-9B-details
Dataset Card for Evaluation run of zelk12/MTM-Merge-gemma-2-9B
Dataset automatically created during the evaluation run of model zelk12/MTM-Merge-gemma-2-9B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/zelk12__MTM-Merge-gemma-2-9B-details.qwen3.5-9B-tau2bench-retail-baseline-traces
Qwen3.5-9B tau2-bench retail baseline traces (n=3 × 114)
Three independent evaluation trials of Qwen3.5-9B (no fine-tune, no memory)
on the full 114-task tau2-bench retail set.
Each trial_N.jsonl is one trial; one JSON object per line, one object per task.
Per-trial pass^1 (canonical reward)
trial_1: 72.8%
trial_2: 69.3%
trial_3: 73.7%
Pooled metrics (n=3 across 114 tasks)
pass^1 = 71.9%
pass^2 = 58.2%
pass^3 = 49.1%
Trace fields (per object)… See the full description on the dataset page: https://huggingface.co/datasets/KermitCO/qwen3.5-9B-tau2bench-retail-baseline-traces.PersonaChat-ERNIE-enhanced-Qwen3.5-9B
Visual Memory Results: personachat-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-9B",
"hf_results_repo": "visual-memory/PersonaChat-ERNIE-enhanced-Qwen3.5-9B",
"results_jsonl": "results/PersonaChat-ERNIE-enhanced-Qwen3.5-9B.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-ERNIE-enhanced-Qwen3.5-9B.ConvAI2-Qwen-original-Qwen3.5-9B
Visual Memory Results: convai2-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-9B",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-original-Qwen3.5-9B",
"results_jsonl": "results/ConvAI2-Qwen-original-Qwen3.5-9B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-original-Qwen3.5-9B.ConvAI2-ERNIE-enhanced-Qwen3.5-9B
Visual Memory Results: convai2-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-9B",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-enhanced-Qwen3.5-9B",
"results_jsonl": "results/ConvAI2-ERNIE-enhanced-Qwen3.5-9B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-enhanced-Qwen3.5-9B.prithivMLmods__GWQ-9B-Preview2-details
Dataset Card for Evaluation run of prithivMLmods/GWQ-9B-Preview2
Dataset automatically created during the evaluation run of model prithivMLmods/GWQ-9B-Preview2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__GWQ-9B-Preview2-details.
