CoolFace
20 results

40b

google /wiki40b Dataset Card for "wiki40b" Dataset Summary Clean-up text for 40+ Wikipedia languages editions of pages correspond to entities. The datasets have train/dev/test splits per language. The dataset is cleaned up by page filtering to remove disambiguation pages, redirect pages, deleted pages, and non-entity pages. Each example contains the wikidata id of the entity, and the full Wikipedia article after page processing that removes non-content sections and structured objects.… See the full description on the dataset page: https://huggingface.co/datasets/google/wiki40b.text10M<n<100M37 likes11k downloads3y agoHugging FaceInfi-MM /InfiMM-WebMath-40B InfiMM-WebMath-40B Dataset ArXiv| PDF This dataset is also discussed in the survey paper A Survey of Deep Learning for Geometry Problem Solving. The accompanying reading list/code for the survey can be found at: https://github.com/majianz/gps-survey InfiMM-WebMath-40B is a large-scale, open-source multimodal dataset specifically designed for mathematical reasoning tasks. It incorporates both text and images, extracted from web documents, to advance the pre-training of Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/Infi-MM/InfiMM-WebMath-40B.textimage-text-to-text10M<n<100M69 likes2.7k downloads1y agoHugging Faceopen-llm-leaderboard-old /details_tiiuae__falcon-40b Dataset Card for Evaluation run of tiiuae/falcon-40b Dataset Summary Dataset automatically created during the evaluation run of model tiiuae/falcon-40b on the Open LLM Leaderboard. The dataset is composed of 124 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_tiiuae__falcon-40b.0 likes674 downloads3y agoHugging FaceLocalDoc /climbmix-40b-az ClimbMix 40B — Azerbaijani A large-scale Azerbaijani text dataset created by translating the English karpathy/climbmix-400b-shuffle dataset into Azerbaijani using Google Translate. Dataset Summary This dataset contains approximately 40 billion tokens of Azerbaijani text, making it one of the largest publicly available Azerbaijani language corpora. It is intended for pretraining and fine-tuning large language models (LLMs) for the Azerbaijani language. Property Value… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/climbmix-40b-az.texttext-generation10M<n<100M2 likes511 downloads6mo agoHugging Faceopen-llm-leaderboard-old /details_AI-Sweden-Models__gpt-sw3-40b Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-40b Dataset Summary Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-40b on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AI-Sweden-Models__gpt-sw3-40b.0 likes311 downloads3y agoHugging Facefn-aka-mur /wiki40b_jaThis dataset is a reformatted version of the Japanese portion of wiki40b dataset. When you use this dataset, please cite the original paper: @inproceedings{guo-etal-2020-wiki, title = "{W}iki-40{B}: Multilingual Language Model Dataset", author = "Guo, Mandy and Dai, Zihang and Vrande{\v{c}}i{\'c}, Denny and Al-Rfou, Rami", booktitle = "Proceedings of the Twelfth Language Resources and Evaluation Conference", month = may, year = "2020", address =… See the full description on the dataset page: https://huggingface.co/datasets/fn-aka-mur/wiki40b_ja.text100K<n<1M5 likes195 downloads3y agoHugging Face