datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Deepthinking-COTDeepthinking-alfworld_and_dbbench_spider_v2
Deepthinking ALFWorld & DBBench Spider v2
AgentBench 評価の 2 タスク(ALFWorld / DBBench)を統合した マルチタスク SFT 訓練データセット。
フォーマット検査・フィルタリング済みの 7,779 件を、サイズ比率に基づく等間隔インターリーブで結合。
Dataset Summary
Metric
Value
Total rows
7,779
ALFWorld
4,884 (62.8%)
DBBench
2,895 (37.2%)
Avg messages per item
18.3
Columns
messages
Interleave method
比率ベース等間隔マージ
Source Datasets
Source
Rows
Description
mark-22/Deepthinking-sft_alfworld_final1
4,884
ALFWorld… See the full description on the dataset page: https://huggingface.co/datasets/mark-22/Deepthinking-alfworld_and_dbbench_spider_v2.Deepthinking-sft_alfworld_final1
Deepthinking SFT ALFWorld (Final)
mark-22/alfworld_combined_shuffled_final(4,884 件)に対し、
GPT-OSS-120B (Groq) を用いて 最初の THOUGHT に「状況分析・常識推論・ステップ分解」を自動挿入 したデータ拡張版。
エージェントが行動前に深く考える(Deep Thinking)能力を SFT で獲得させることを目的としている。
Dataset Summary
Metric
Value
Total rows
4,884
Source
mark-22/alfworld_combined_shuffled_final
Augmentation model
GPT-OSS-120B (via Groq API)
Avg messages per item
23.0
Items with THOUGHT + ACTION
4,884 / 4,884 (100%)
Columns
original_id… See the full description on the dataset page: https://huggingface.co/datasets/mark-22/Deepthinking-sft_alfworld_final1.Deepthinking-alfworld_and_dbbench_spider_v1Deepthinking-sft_alfworld_v4Deepthinking-sft_alfworld_test2
