mon
Datasets
All datasets matching “mon”monet
Dataset Card for MONET
MONET (Massive, Open, Non-redundant and Enriched Text-to-image dataset) is a large-scale, curated image-text dataset designed for training text-to-image (T2I) systems. It contains 103.8 million high-quality image-text pairs distilled from 2.9 billion raw pairs across nine heterogeneous open sources (6 real and 3 synthetic) through successive stages of safety filtering, domain-based filtering, exact and near-duplicate removal, and re-captioning with… See the full description on the dataset page: https://huggingface.co/datasets/jasperai/monet.pile-uncopyrighted
Pile Uncopyrighted
In response to authors demanding that LLMs stop using their works, here's a copy of The Pile with all copyrighted content removed.Please consider using this dataset to train your future LLMs, to respect authors and abide by copyright law.Creating an uncopyrighted version of a larger dataset (ie RedPajama) is planned, with no ETA.
MethodologyCleaning was performed by removing everything from the Books3, BookCorpus2, OpenSubtitles, YTSubtitles, and OWT2… See the full description on the dataset page: https://huggingface.co/datasets/monology/pile-uncopyrighted.montok
MonTok: A Suite of Monolingual Tokenizers
This is a set of monolingual tokenizers for 98 languages. For each language, there are Unigram, BPE, and SuperBPE tokenizers, ranging in vocabulary size from around 6k to over 200k.
Training Details
Training Data
All tokenizers are trained on samples of the data used to the train the Goldfish language models.
The tokenizers were either trained on scaled or unscaled data. This refers to whether the models are trained on… See the full description on the dataset page: https://huggingface.co/datasets/catherinearnett/montok.hermes-system-monitor-v2VLN_Dataset_2This repository contains encrypted visual features for an ongoing academic research project. Decryption keys are managed internally for reproducibility.
Charge-040_0040-Sparse-Mono
