datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mala-monolingual-integration
MaLA Corpus: Massive Language Adaptation Corpus
This is the noisy version that integrates texts from different sources.
Dataset Summary
The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-integration.Integration_Problem_Set
Dataset Project: Integration Problem Set
Author: Ruoxin Wang
Date: Nov.26.2024
GitHub Repository
The code is available at https://github.com/WendyWAAAAANG/Integration_Problem_Set
Executive Summary
This project aims to create a dataset specifically designed for the automatic generation of integration problems. It will serve as the basis for training LLMs to generate integration problems. The problems will be labeled by difficulty and… See the full description on the dataset page: https://huggingface.co/datasets/Roxanne-WANG/Integration_Problem_Set.
