datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cleand_moremilk_ToT-Biology元データ: https://huggingface.co/datasets/moremilk/ToT-Biology
データ件数: 5,752
平均トークン数: 675
最大トークン数: 1,105
合計トークン数: 3,881,334
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 19.3 MB
加工内容:
長文フィルタリング: トークナイズ処理の負荷を軽減するため、事前に文字列が極端に長い行を除外します。
question 列: 6,000文字を超える行を除外。
metadata 列: 80,000文字を超える行を除外。
metadata フィールドの展開:
metadata 列に含まれるJSON形式のデータから reasoning と difficulty の値を抽出します。
reasoning は thought という新しい列に格納します。
difficulty は difficulty という新しい列に格納します。
処理後、元の metadata 列は削除されます。
繰り返し表現の除去:
thought… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_moremilk_ToT-Biology.tot-cwq-plan-sft-outputs34-rule-full-pw4-expand-labels-v2
ToT CWQ Plan SFT - outputs34_rule_full_pw4_expand_labels_v2
Merged SFT output from local run outputs34_rule_full_pw4_expand_labels_v2.
Version ID
local output dir: tot/sft/outputs34_rule_full_pw4_expand_labels_v2
file: cwq_train_plan.no_mid.jsonl
dataset: CWQ
grouping backend: TOT_REL_GROUPING_BACKEND=rules
parallel workers: 4
strict expand parity: enabled
nested expand labels: enabled
Main difference from earlier runs
This version renders nested Expand… See the full description on the dataset page: https://huggingface.co/datasets/YF0808/tot-cwq-plan-sft-outputs34-rule-full-pw4-expand-labels-v2.
