combine
Datasets
All datasets matching “combine”mosaic-combine-all
Mosaic format for combine all dataset to train Malaysian LLM
This repository is to store dataset shards using mosaic format.
prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-all.ipynb
using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer
4096 context length.
how-to
git clone,
git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-combine-all
load it,
from streaming import LocalDataset
import numpy… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-combine-all.so-combined-engThis dataset was created using LeRobot.
Dataset Description
The English version of this dataset integrates 598 open-source community datasets into a single unified corpus, comprising 22,709 episodes and approximately 9.4 million frames across 563 distinct tasks. Several transformations were applied to ensure standardization and data quality:
Camera view normalizationBecause community datasets do not follow a consistent naming scheme for camera viewpoints, we used the… See the full description on the dataset page: https://huggingface.co/datasets/dunnolab/so-combined-eng.v0-train-combinedso-combined-ruДатасет создан при помощи библиотеки LeRobot.
Описание датасета
Русскоязычная версия данного датасета объединяет 598 открытых датасетов сообщества в единый унифицированный корпус, включающий 22 709 эпизодов и примерно 9,4 миллиона кадров по 563 различным задачам. Для обеспечения стандартизации и качества данных были выполнены следующие преобразования:
Нормализация ракурсов камеры
Поскольку датасеты сообщества не используют общепринятую схему именования ракурсов… See the full description on the dataset page: https://huggingface.co/datasets/dunnolab/so-combined-ru.tts-dataset-combinedcommit0_combined
