datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
marianne_pdf_7marianne_pdf_9marianne_pdf_5marianne_pdf_3marianne_pdf_8marianne_pdf_4marianne_pdf_10hal_pdf_extraxares_llm_datafinetune_data
tdro-llm/finetune_data
tDRO: Task-level Distributionally Robust Optimization for Large Language Model-based Dense Retrieval. Guangyuan Ma, Yongliang Ma, Xing Wu, Zhenpeng Su, Ming Zhou and Songlin Hu.
This repo contains all fine-tuning data for Large Language Model-based Dense Retrieval. Please refer to this repo for details to reproduce.
A total of 25 heterogeneous retrieval fine-tuning datasets with Hard Negatives and Deduplication (with test sets) are listed as belows.… See the full description on the dataset page: https://huggingface.co/datasets/tdro-llm/finetune_data.JP-LLM-Corpus-PII-Filtered-10B
CommonCrawl Japanese (Filtered PPI) Dataset
本データセットは、CommonCrawlより抽出した約100億(10B)トークン規模の日本語テキストデータから、特に配慮が必要な「要配慮個人情報」をフィルタリング処理したものです。
データセットの概要
元データソース: CommonCrawl(https://commoncrawl.org/)
トークン数: 約10Bトークン
言語: 日本語
処理内容: 要配慮個人情報をルールベースおよび機械学習分類器を用いてフィルタリング
フィルタリングには以下のコードを使用しております。https://github.com/matsuolab/jp-llm-corpus-pii-filter/
注意事項
本データセットは、非常に大規模なテキストから自動的に要配慮個人情報を除去したものであり、完全な排除を保証するものではありません。そのため、二次的な活用に際しては、目的に応じた適切な管理・配慮が必要です。… See the full description on the dataset page: https://huggingface.co/datasets/matsuo-lab/JP-LLM-Corpus-PII-Filtered-10B.LLMMultiLabelefficient_llm
Data V4 for NeurIPS LLM Challenge
Contains 70949 samples collected from Huggingface:
Math: 1273
gsm8k
math_qa
math-eval/TAL-SCQ5K
TAL-SCQ5K-EN
meta-math/MetaMathQA
TIGER-Lab/MathInstruct
Science: 42513
lighteval/mmlu - 'all', "split": 'auxiliary_train'
lighteval/bbq_helm - 'all'
openbookqa - 'main'
ComplexQA: 2940
ARC-Challenge
ARC-Easy
piqa
social_i_qa
Muennighoff/babi
Rowan/hellaswag
ComplexQA1: 2060
medmcqa
winogrande_xl,
winogrande_debiased
boolq
sciq
CNN: 2787… See the full description on the dataset page: https://huggingface.co/datasets/transZ/efficient_llm.LCB128_Llama3.1-8B-Inst_LLM-as-judge_256_32create_llm_from_scratchLLM_SAFETY_CHECK_OOD
