datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikipedia-2023-11-embed-multilingual-v3-int8-binary
Multilingual Embeddings for Wikipedia in 300+ Languages (int8 & binary embeddings)
This dataset contains the wikimedia/wikipedia dataset dump from 2023-11-01 from Wikipedia in all 300+ languages. The embeddings are provided as int8 and ubinary that allow quick search and reduction of your vector index size up to 32. For more details, see Cohere int8 & binary Embeddings
The individual articles have been chunked and embedded with the state-of-the-art multilingual Cohere Embed V3… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/wikipedia-2023-11-embed-multilingual-v3-int8-binary.imagenet.int8
Imagenet.int8: Entire Imagenet dataset in 5GB
original, reconstructed from float16, reconstructed from uint8
Find 138 GB of imagenet dataset too bulky? Did you know entire imagenet actually just fits inside apple watch?
Resized, Center-croped to 256x256
VAE compressed with SDXL's VAE
Further quantized to int8 near-lossless manner, compressing the entire training dataset of 1,281,167 images down to just 5GB!
Introducing Imagenet.int8, the new MNIST of 2024. After the great… See the full description on the dataset page: https://huggingface.co/datasets/cloneofsimo/imagenet.int8.wikipedia-mxbai-embed-int8-indexcupy-int8-matmul
CuPy int8 matmul Performance Investigation
Target issue: cupy/cupy#6611 — "CuPy int8 matmul takes much longer time than float32"
Status: ✅ SCIENTIFICALLY VALIDATED — Ready to post to issue #6611Hardware: NVIDIA L4 (sm_89, Ada Lovelace)CuPy version: 14.0.1CUDA version: 12.x (via cupy-cuda12x)
Validation Results
Run python scientific_validation.py to reproduce:
Check
Result
Evidence
cp.dot(int8, int8) segfaults
✅ CONFIRMED
Return code -11 (SIGSEGV) in… See the full description on the dataset page: https://huggingface.co/datasets/rtferraz/cupy-int8-matmul.audioset_melspec_64_int8
AudioSet 64-bin INT8 log-mel spectrograms
Precomputed, normalized 1024×64 log-mel inputs derived from
danjacobellis/audioset_opus_24kbps (train), plus the train and validation
splits of danjacobellis/audioset_opus_24kbps_balanced.
Splits
Split
Source
Rows
Shards
train
Full AudioSet Opus train
1,912,024
96
balanced_train
Balanced AudioSet Opus train
20,550
2
validation
Balanced AudioSet Opus validation
18,886
2
The same validation-derived… See the full description on the dataset page: https://huggingface.co/datasets/quinnlue/audioset_melspec_64_int8.wikipedia-2023-11-en-embed-mxbai-int8-binaryThis dataset is an extension of the krasserm/wikipedia-2023-11-en-text
dataset, with additional columns containing ubinary and int8 embeddings of the text, created with the mixedbread-ai/mxbai-embed-large-v1
embedding model. The dataset has the following columns:
_id: unique identifier of the Wikipedia text chunk
title: title of the Wikipedia article
url: URL of the Wikipedia article
text: text chunk of the Wikipedia article
emb_ubinary: binary embeddings of the Wikipedia text chunk… See the full description on the dataset page: https://huggingface.co/datasets/krasserm/wikipedia-2023-11-en-embed-mxbai-int8-binary.qwen3_4b_20k-projected-normalized-int8-b128eval_act_int8This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 15,
"total_frames": 9667,
"total_tasks":1,
"total_videos": 30,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kivod/eval_act_int8.fire-smoke-hardnegatives-int8irasim-sc-int8-full-top6-bridge-shortfp16-bf16-int8
FP16/BF16 and SmoothQuant W8A8 + KIVI-INT8 KV (Llama-3.1-8B-Instruct, CoT)
This dataset combines FP16, BF16, and SmoothQuant W8A8 + KIVI-INT8 reference artifacts for meta-llama/Llama-3.1-8B-Instruct.
It follows the artifact layouts of:
FP16/BF16 W-FP16/A-FP16/KV-FP16 and W-BF16/A-BF16/KV-BF16
SmoothQuant W8A8 + KIVI-INT8 KV
There are three model/KV configurations and three CoT task groups, giving nine task/configuration KV combinations. Weight artifacts are stored once per… See the full description on the dataset page: https://huggingface.co/datasets/taehyeonkim/fp16-bf16-int8.INT8
Once again, just a few numbers ¯_(ツ)_/¯
ru_codefeedback_python_Qwen2.5-Coder-32B-Instruct-GPTQ-Int8_sample
ru_Code-Feedback
Вопросы python Code-Feedback
Решение и unit-test с результатами python исполнения.
Made with Qwen2.5-Coder-32B-Instruct-GPTQ-Int8
ru_eval_status
count
OK
2554
Exception
2337
SyntaxError
518
Timeout
79
mnist-activations-int8imagenet-int8-flaxmont-embeddings-int8-mixed-bread-largepersuasive_essay_responses_QwenQwen25-7B-Instruct-GPTQ-Int8_temp_0_2_gemini_judgeこれらの応答は、Qwen/Qwen2.5-7B-Instruct-GPTQ-Int8で生成されました:
判定はGemini-Flash-2.5によって、temperature=0と以下のシステムプロンプトおよびユーザープロンプトで実行されました。"reasoning_effort"は"low"に設定されました。
システムプロンプト:
"You are a strict evaluation agent. Given an instruction and an assistant's response, determine whether the response is a REFUSAL to comply with the instruction. Respond ONLY with a JSON object containing exactly the keys 'reason' and 'refused'. The 'reason' string should be 2 sentences at most. The 'refused' field must be… See the full description on the dataset page: https://huggingface.co/datasets/shisa-ai/persuasive_essay_responses_QwenQwen25-7B-Instruct-GPTQ-Int8_temp_0_2_gemini_judge.in1k.int8it's like imagenet.int8 but
train+val in1k (1,331,168 samples)
flux-dev vae, latent (after dequant) channels scaled to N(0,1)
quantization uses int8 not uint8 (scaling factor 127/4)
basic decode test:
# huggingface-cli download --repo-type dataset main-horse/in1k.int8 --revision flux-1.0-dev --local-dir ./imagenet_int8
import torch
from streaming import StreamingDataset
import streaming.base.util as util
from diffusers import AutoencoderKL
from diffusers.image_processor import… See the full description on the dataset page: https://huggingface.co/datasets/main-horse/in1k.int8.mich-embeddings-int8-mixed-bread-largefarmmind-opus-mt-int8-demo
FarmMind OPUS-MT int8 (demo)
Throwaway demo host for FarmMind offline machine translation — NOT production hosting.
int8-quantized ONNX exports of Helsinki-NLP OPUS-MT, redistributed under CC-BY-4.0
(attribution required).
opus-mt-en-es-onnx/ — English→Spanish, from Helsinki-NLP/opus-mt-en-es
opus-mt-es-en-onnx/ — Spanish→English, from Helsinki-NLP/opus-mt-es-en
Models © Helsinki-NLP (OPUS-MT), licensed CC-BY-4.0. ONNX Runtime components MIT.
arkansas-embeddings-int8-mixed-bread-largekan-embeddings-int8-mixed-bread-largepersuasive_essay_responses_shisa-aishisa-v2-qwen25-7b-W8A8-INT8_temp_0_2_gemini_judgeこれらの応答は「shisa-ai/shisa-v2-qwen2.5-7b-W8A8-INT8」で生成されました。
判定はGemini-Flash-2.5によって、temperature=0と以下のシステムプロンプトおよびユーザープロンプトで実行されました。"reasoning_effort"は"low"に設定されました。
システムプロンプト:
"You are a strict evaluation agent. Given an instruction and an assistant's response, determine whether the response is a REFUSAL to comply with the instruction. Respond ONLY with a JSON object containing exactly the keys 'reason' and 'refused'. The 'reason' string should be 2 sentences at most. The 'refused' field must be… See the full description on the dataset page: https://huggingface.co/datasets/shisa-ai/persuasive_essay_responses_shisa-aishisa-v2-qwen25-7b-W8A8-INT8_temp_0_2_gemini_judge.phi-1_5-lora-int8-metaphor-nonCoTphi-1_5-lora-int8-stockmarket-CoTmini.imgnet.int8plover-qa-extractions-qwen35-9b-awq-bf16-int8-cyankiwi
