datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cleand_microsoft_rStar-Coder元データ: https://huggingface.co/datasets/microsoft/rStar-Coder
データ件数: 269,863
平均トークン数: 11674
最大トークン数: 31,184
合計トークン数: 3,150,447,484
ファイル形式: JSONL
ファイルサイズ: 不明
加工内容
synthetic_sftを使用
トークン処理が重たいので、文字数でフィルター
seed_question < 6000
generation < 80000
thinkタグ除去 が中途半端なものを除外
トークナイズ処理(速度向上アップデート
繰り返し除去
rstar_ppm
Dataset for rStar-Math
https://github.com/microsoft/rStar
gsm8k_qwen_llama_rstargsm8khard_llama_qwen_rstarhuggingface_gms8k_rstar_newwhat-is-art-rst
RST-Inspired Rhetorical Annotation of Tolstoy's What Is Art?
A discourse annotation dataset for Leo Tolstoy's What Is Art? (1904), produced using a functionally adapted RST scheme designed for argumentative and philosophical prose. The annotations were created as part of a study on stance-conditioned fallacy judgment in philosophical argumentation, currently under review.
Source Text
Tolstoy, L. (1904). What is art? Funk & Wagnalls. Retrieved from Project… See the full description on the dataset page: https://huggingface.co/datasets/shoochoon/what-is-art-rst.bunnycore__Phi-4-RStock-v0.1-details
Dataset Card for Evaluation run of bunnycore/Phi-4-RStock-v0.1
Dataset automatically created during the evaluation run of model bunnycore/Phi-4-RStock-v0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__Phi-4-RStock-v0.1-details.gsm8k_rstar_dpo_Datasetmath_qwen_gemma_qwen_rstarSTG_llama_qwen_rstarMATH_llama_qwen_rstarhf_gsm8k_rstar_small
