babe
Datasets
All datasets matching “babe”piqaBabelDOC-Assets
BabelDOC-Assets
Font and other resource files relied on by BabelDOC and pdf2zh
BabelDOC is a PDF translation library, pdf2zh is a PDF translation tool.
Fonts and Licenses
Go Noto Universal: THE UNLICENSE
Pal3love/Source-Han-TrueType: SIL OPEN FONT LICENSE Version 1.1
lxgw/LxgwWenKaiGB: OFL-1.1 License
lxgw/LxgwWenkaiTC: OFL-1.1 License
fontworks-fonts/Klee: OFL-1.1 License
fonts-archive/MaruBuri: License
Noto Serif/Noto Sans: SIL OPEN FONT LICENSE Version 1.1… See the full description on the dataset page: https://huggingface.co/datasets/awwaawwa/BabelDOC-Assets.paul_graham_essaysmultilingual_mmluMMLU professionally translated into 14 languages using professional human translators, sourced from OpenAI's simple-eval.
Original files:
english: https://openaipublic.blob.core.windows.net/simple-evals/mmlu.csv
multilingual: https://openaipublic.blob.core.windows.net/simple-evals/mmlu_{language}.csv where language one of "AR-XY", "BN-BD", "DE-DE", "ES-LA", "FR-FR", "HI-IN", "ID-ID", "IT-IT", "JA-JP", "KO-KR", "PT-BR", "ZH-CN", "SW-KE", "YO-NG", "EN-US"
global-piqa-evalslogiqa2The dataset is an amendment and re-annotation of LogiQA in 2020, a large-scale logical reasoning reading comprehension dataset adapted from the Chinese Civil Service Examination. We increase the data size, refine the texts with manual translation by professionals, and improve the quality by removing items with distinctive cultural features like Chinese idioms. Furthermore, we conduct a fine-grained annotation on the dataset and turn it into a two-way natural language inference (NLI) task, resulting in 35k premise-hypothesis pairs with gold labels, making it the first large-scale NLI dataset for complex logical reasoning
