deep thinking
Qwen3-30B-A3B-Thinking-2507-Deepseek-v3.1-Distill-FP32-GGUFgemma-3-12b-it-vl-Deepseek-v3.1-Heretic-Uncensored-Thinking-i1-GGUFQwen3.5-13B-GLM-4.7-Flash-DeepSeek-Polaris-Grande-Deep-Thinking-i1-GGUFQwen3.5-2B-GPT-5.1-HighIQ-Deep-Thinking-i1-GGUFMistral-Nemo-2407-Instruct-12B-Deep-Thinking-Claude-Gemini-GPT5.2-i1-GGUFgemma-3-16b-it-BIG-G-GLM4.7-Flash-Valhalla-Heretic-Uncensored-Deep-Thinking-i1-GGUFQwen3.5-9B-DeepSeek-3.2-Intense-Auto-Variable-Thinking-i1-GGUFQwen3-30B-A3B-Thinking-2507-Deepseek-v3.1-Distill-V2-FP32-i1-GGUF
details_xDAN-AI__xDAN-L1Mix-DeepThinking-v2
Dataset Card for Evaluation run of xDAN-AI/xDAN-L1Mix-DeepThinking-v2
Dataset automatically created during the evaluation run of model xDAN-AI/xDAN-L1Mix-DeepThinking-v2 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_xDAN-AI__xDAN-L1Mix-DeepThinking-v2.Deepthinking-COTDeepthinking-alfworld_and_dbbench_spider_v2
Deepthinking ALFWorld & DBBench Spider v2
AgentBench 評価の 2 タスク(ALFWorld / DBBench)を統合した マルチタスク SFT 訓練データセット。
フォーマット検査・フィルタリング済みの 7,779 件を、サイズ比率に基づく等間隔インターリーブで結合。
Dataset Summary
Metric
Value
Total rows
7,779
ALFWorld
4,884 (62.8%)
DBBench
2,895 (37.2%)
Avg messages per item
18.3
Columns
messages
Interleave method
比率ベース等間隔マージ
Source Datasets
Source
Rows
Description
mark-22/Deepthinking-sft_alfworld_final1
4,884
ALFWorld… See the full description on the dataset page: https://huggingface.co/datasets/mark-22/Deepthinking-alfworld_and_dbbench_spider_v2.Deepthinking-sft_alfworld_final1
Deepthinking SFT ALFWorld (Final)
mark-22/alfworld_combined_shuffled_final(4,884 件)に対し、
GPT-OSS-120B (Groq) を用いて 最初の THOUGHT に「状況分析・常識推論・ステップ分解」を自動挿入 したデータ拡張版。
エージェントが行動前に深く考える(Deep Thinking)能力を SFT で獲得させることを目的としている。
Dataset Summary
Metric
Value
Total rows
4,884
Source
mark-22/alfworld_combined_shuffled_final
Augmentation model
GPT-OSS-120B (via Groq API)
Avg messages per item
23.0
Items with THOUGHT + ACTION
4,884 / 4,884 (100%)
Columns
original_id… See the full description on the dataset page: https://huggingface.co/datasets/mark-22/Deepthinking-sft_alfworld_final1.Deepthinking-alfworld_and_dbbench_spider_v1Deepthinking-sft_alfworld_v4
