datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
details_matsuo-lab__weblab-10b
Dataset Card for Evaluation run of matsuo-lab/weblab-10b
Dataset Summary
Dataset automatically created during the evaluation run of model matsuo-lab/weblab-10b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_matsuo-lab__weblab-10b.JP-LLM-Corpus-PII-Filtered-10B
CommonCrawl Japanese (Filtered PPI) Dataset
本データセットは、CommonCrawlより抽出した約100億(10B)トークン規模の日本語テキストデータから、特に配慮が必要な「要配慮個人情報」をフィルタリング処理したものです。
データセットの概要
元データソース: CommonCrawl(https://commoncrawl.org/)
トークン数: 約10Bトークン
言語: 日本語
処理内容: 要配慮個人情報をルールベースおよび機械学習分類器を用いてフィルタリング
フィルタリングには以下のコードを使用しております。https://github.com/matsuolab/jp-llm-corpus-pii-filter/
注意事項
本データセットは、非常に大規模なテキストから自動的に要配慮個人情報を除去したものであり、完全な排除を保証するものではありません。そのため、二次的な活用に際しては、目的に応じた適切な管理・配慮が必要です。… See the full description on the dataset page: https://huggingface.co/datasets/matsuo-lab/JP-LLM-Corpus-PII-Filtered-10B.details_matsuo-lab__weblab-10b-instruction-sft
Dataset Card for Evaluation run of matsuo-lab/weblab-10b-instruction-sft
Dataset Summary
Dataset automatically created during the evaluation run of model matsuo-lab/weblab-10b-instruction-sft on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_matsuo-lab__weblab-10b-instruction-sft.
