int8
Datasets
All datasets matching “int8”wikipedia-2023-11-embed-multilingual-v3-int8-binary
Multilingual Embeddings for Wikipedia in 300+ Languages (int8 & binary embeddings)
This dataset contains the wikimedia/wikipedia dataset dump from 2023-11-01 from Wikipedia in all 300+ languages. The embeddings are provided as int8 and ubinary that allow quick search and reduction of your vector index size up to 32. For more details, see Cohere int8 & binary Embeddings
The individual articles have been chunked and embedded with the state-of-the-art multilingual Cohere Embed V3… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/wikipedia-2023-11-embed-multilingual-v3-int8-binary.imagenet.int8
Imagenet.int8: Entire Imagenet dataset in 5GB
original, reconstructed from float16, reconstructed from uint8
Find 138 GB of imagenet dataset too bulky? Did you know entire imagenet actually just fits inside apple watch?
Resized, Center-croped to 256x256
VAE compressed with SDXL's VAE
Further quantized to int8 near-lossless manner, compressing the entire training dataset of 1,281,167 images down to just 5GB!
Introducing Imagenet.int8, the new MNIST of 2024. After the great… See the full description on the dataset page: https://huggingface.co/datasets/cloneofsimo/imagenet.int8.wikipedia-mxbai-embed-int8-indexcupy-int8-matmul
CuPy int8 matmul Performance Investigation
Target issue: cupy/cupy#6611 — "CuPy int8 matmul takes much longer time than float32"
Status: ✅ SCIENTIFICALLY VALIDATED — Ready to post to issue #6611Hardware: NVIDIA L4 (sm_89, Ada Lovelace)CuPy version: 14.0.1CUDA version: 12.x (via cupy-cuda12x)
Validation Results
Run python scientific_validation.py to reproduce:
Check
Result
Evidence
cp.dot(int8, int8) segfaults
✅ CONFIRMED
Return code -11 (SIGSEGV) in… See the full description on the dataset page: https://huggingface.co/datasets/rtferraz/cupy-int8-matmul.audioset_melspec_64_int8
AudioSet 64-bin INT8 log-mel spectrograms
Precomputed, normalized 1024×64 log-mel inputs derived from
danjacobellis/audioset_opus_24kbps (train), plus the train and validation
splits of danjacobellis/audioset_opus_24kbps_balanced.
Splits
Split
Source
Rows
Shards
train
Full AudioSet Opus train
1,912,024
96
balanced_train
Balanced AudioSet Opus train
20,550
2
validation
Balanced AudioSet Opus validation
18,886
2
The same validation-derived… See the full description on the dataset page: https://huggingface.co/datasets/quinnlue/audioset_melspec_64_int8.wikipedia-2023-11-en-embed-mxbai-int8-binaryThis dataset is an extension of the krasserm/wikipedia-2023-11-en-text
dataset, with additional columns containing ubinary and int8 embeddings of the text, created with the mixedbread-ai/mxbai-embed-large-v1
embedding model. The dataset has the following columns:
_id: unique identifier of the Wikipedia text chunk
title: title of the Wikipedia article
url: URL of the Wikipedia article
text: text chunk of the Wikipedia article
emb_ubinary: binary embeddings of the Wikipedia text chunk… See the full description on the dataset page: https://huggingface.co/datasets/krasserm/wikipedia-2023-11-en-embed-mxbai-int8-binary.
