datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tinycompany__SigmaBoi-base-details
Dataset Card for Evaluation run of tinycompany/SigmaBoi-base
Dataset automatically created during the evaluation run of model tinycompany/SigmaBoi-base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tinycompany__SigmaBoi-base-details.tinycompany__SigmaBoi-nomic1.5-fp32-details
Dataset Card for Evaluation run of tinycompany/SigmaBoi-nomic1.5-fp32
Dataset automatically created during the evaluation run of model tinycompany/SigmaBoi-nomic1.5-fp32
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tinycompany__SigmaBoi-nomic1.5-fp32-details.creative_writing
Creative Writing & Metrics Evaluation Dataset
Dataset Description
Each row is one human-written continuation of a creative-writing prompt, scored automatically by four LLM judges (gemini-2.0-flash, gemini-3.8-flash, gpt-4o, gpt-5.6-terra) and a set of traditional NLP metrics, and reviewed independently by multiple human raters on the same criteria.
The dataset consists of responses to creative writing prompts. Each prompt specifically contained a direction to… See the full description on the dataset page: https://huggingface.co/datasets/sigma-ai-research/creative_writing.tinycompany__SigmaBoi-nomic-moe-details
Dataset Card for Evaluation run of tinycompany/SigmaBoi-nomic-moe
Dataset automatically created during the evaluation run of model tinycompany/SigmaBoi-nomic-moe
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tinycompany__SigmaBoi-nomic-moe-details.tinycompany__SigmaBoi-ib-details
Dataset Card for Evaluation run of tinycompany/SigmaBoi-ib
Dataset automatically created during the evaluation run of model tinycompany/SigmaBoi-ib
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tinycompany__SigmaBoi-ib-details.tinycompany__SigmaBoi-bgem3-details
Dataset Card for Evaluation run of tinycompany/SigmaBoi-bgem3
Dataset automatically created during the evaluation run of model tinycompany/SigmaBoi-bgem3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tinycompany__SigmaBoi-bgem3-details.cleand_cw18_lean-six-sigma-cot-500元データ: https://huggingface.co/datasets/cw18/lean-six-sigma-cot-500
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/lean-six-sigma-cot-500
データ件数: 215
平均トークン数: 514
最大トークン数: 591
合計トークン数: 110,520
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 602.9 KB
加工内容:
文字列長によるフィルタリング:
instruction列(質問)の文字数が6000文字を超える行を除外しました。
output列(思考)の文字数が80000文字を超える行を除外しました。
思考タグの除去と分割:
IS_THINKTAGがFalseに設定されているため、output列をSPLIT_KEYWORD (**Final Toolset… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_cw18_lean-six-sigma-cot-500.tinycompany__SigmaBoi-bge-m3-details
Dataset Card for Evaluation run of tinycompany/SigmaBoi-bge-m3
Dataset automatically created during the evaluation run of model tinycompany/SigmaBoi-bge-m3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tinycompany__SigmaBoi-bge-m3-details.tinycompany__SigmaBoi-nomic1.5-details
Dataset Card for Evaluation run of tinycompany/SigmaBoi-nomic1.5
Dataset automatically created during the evaluation run of model tinycompany/SigmaBoi-nomic1.5
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tinycompany__SigmaBoi-nomic1.5-details.
