datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hypergraph_openthoughts30k
Hypergraph OpenThoughts Math 30K
Reasoning hypergraphs generated for the 29,434 examples in
siyanzhao/Openthoughts_math_30k_opsd.
Generation
Model: Qwen/Qwen3.6-35B-A3B-FP8
Thinking mode: disabled
Construction: semantic-step segmentation followed by primary-support DAG induction
Graph constraint: at most one earlier-step parent per semantic step
Processing order: source dataset row order
Schema
Each JSONL record contains:
row_index: source… See the full description on the dataset page: https://huggingface.co/datasets/dvtiendat/hypergraph_openthoughts30k.open-thoughts__OpenThinker-7B-details
Dataset Card for Evaluation run of open-thoughts/OpenThinker-7B
Dataset automatically created during the evaluation run of model open-thoughts/OpenThinker-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/open-thoughts__OpenThinker-7B-details.cleand_openthought312_dif9_tiny元データ: https://huggingface.co/datasets/LLMTeamAkiyama/clean_openthought312_difficulty_9_filterd
データ件数: 1,456
平均トークン数: 5,894
最大トークン数: 8,186
合計トークン数: 8,581,562
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 33.2 MB
加工内容:
元データに対して、token数を8912以下に制限したテスト用tiny版
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/openthoughts3/clean_openthoughts3_tiny_pickup.ipynb
clean_openthought312_difficulty_9_filterd元データ: https://huggingface.co/datasets/open-thoughts/OpenThoughts3-1.2M
diffculty 9でさらにフィルタリングしたもの
データ件数: 14,339
平均トークン数: 13370
最大トークン数: 16,808
合計トークン数: 191,708,678
ファイル形式: JSONL
ファイルサイズ: 723.9 MB
bunnycore__Llama-3.2-3B-Bespoke-Thought-details
Dataset Card for Evaluation run of bunnycore/Llama-3.2-3B-Bespoke-Thought
Dataset automatically created during the evaluation run of model bunnycore/Llama-3.2-3B-Bespoke-Thought
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__Llama-3.2-3B-Bespoke-Thought-details.clean_openthought312_difficulty_9_qwentoken元データ: https://huggingface.co/datasets/LLMTeamAkiyama/clean_openthought312_difficulty_9_filterd
データ件数: 14,339
平均トークン数: 13,367
最大トークン数: 16,805
合計トークン数: 191,665,652
ファイル形式: JSONL
ファイル分割数: 3
合計ファイルサイズ: 724.7 MB
加工内容:
**tokenizeをQwen235B-A22Bで再度トークン化したものを出力
使用したコード
https://github.com/LLMTeamAkiyama/0-data_prepare/blob/master/src/openthoughts3/clean_openthoughts3_9_qwentoken.ipynb
Solshine__Llama-3-1-big-thoughtful-passthrough-merge-2-details
Dataset Card for Evaluation run of Solshine/Llama-3-1-big-thoughtful-passthrough-merge-2
Dataset automatically created during the evaluation run of model Solshine/Llama-3-1-big-thoughtful-passthrough-merge-2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Solshine__Llama-3-1-big-thoughtful-passthrough-merge-2-details.OpenThoughts2-1M-fiilter
