datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imatrix-calibration
Importance Matrix Calibration Datasets
This repository provides calibration datasets used to generate importance matrices (imatrix), which are required to minimize errors when quantizing models with LLaMA C++.
The llama-imatrix program cannot handle parquet files directly and thus requires them to be converted into text format first. There are many ways to do this but a simple approach is to use DuckDB with the following command: duckdb -noheader -ascii -c "SELECT content FROM… See the full description on the dataset page: https://huggingface.co/datasets/eaddario/imatrix-calibration.imatrix-storageimatrixQwQ-32B-abliterated-131k-GGUF-Yarn-Imatrix
QwQ-32B-Abliterated-131k-GGUF-Yarn-Imatrix
High-Fidelity Semantic Simulation & Orchestration AI Model
Will this pass the random stupid benchmarks that exist today? I don't know, nor care. I don't need my local AI model to know some random city capital of a foreign country. I need a local AI model that can simulate with high semantic fidelity. Why? Because your AI may be able to spit random facts. I want an AI that knows when to Google facts. I want an AI that tracks hundreds of… See the full description on the dataset page: https://huggingface.co/datasets/magiccodingman/QwQ-32B-abliterated-131k-GGUF-Yarn-Imatrix.imatrix
Input files for generating the Importance Matrix
Which file to use for generating the importance matrix
Not all importance matrices are equal. The best results are obtained when using a source file similar to the
training data. Size also matters: the bigger the model (eg: 70b vs 13b) and the higher the quant (eg: q6k_ vs iq3_xs),
the bigger the source file needs to be to make an impact. Multiple input files can be combined if needed;
for example:
cat multilingual.txt… See the full description on the dataset page: https://huggingface.co/datasets/froggeric/imatrix.imatrix-from-wiki-trainThis repository contains importance matrix datasets for use with the improved quantization methods recently added to llama.cpp.
The importance matrix has been computed using wiki.train.raw as training data.
Hope the file names are self-explanatory.
To use, after cloning this repo, for e.g. Mixtral-8x7B and Q4_K_M quantization, use
./quantize --imatrix path_to_repo/mixtral-8x7b.imatrix path_to_model ggml-model-q4k-m.gguf Q4_K_M
bartowski-imatrix-v5-semantic
Bartowski iMatrix Calibration v5 (Semantic Chunking)
A processed version of bartowski's v5 imatrix calibration data using semantic boundary detection optimized for the v5 data structure.
Dataset Summary
Metric
Value
Total samples
2,075
Chunking method
V5-optimized semantic boundary detection
Chunk size
200+ characters (no upper limit, preserves document integrity)
Languages
English, German, Spanish, French, Italian, Swedish, Russian, Arabic, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/bartowski-imatrix-v5-semantic.korean-imatrix-calibration-corpus
Korean imatrix Calibration Corpus — KO-i1 보정 코퍼스
한국어 중심 imatrix 보정 코퍼스의 첫 공개 릴리스 (우리가 아는 한).
공개 GGUF 양자화 생태계의 importance matrix는 거의 전부 영어 위주 코퍼스로
수집됩니다. 그 결과 한국어 토큰 분포에서의 양자화 손실이 체계적으로 커집니다.
이 데이터셋은 그 공백을 메우기 위해 만들어졌고, 실측으로 효과가 입증됐습니다.
실측 효과 (이 코퍼스로 만든 KO-i1 릴리스들)
릴리스
비교 대상
결과
kanana-1.5-8b KO-i1
영어 보정 i1
저비트 KLD -5~6% (IQ2_M 3.3σ), 비트 낮을수록 이득 증가
Qwen3.6-35B-A3B KO-i1
영어 보정 i1
전 타입 우세, -5.1~-6.8% (최대 4.3σ), MoE는 4비트도 유의
Qwen3.8-27B-abl KO-i1
정적 양자… See the full description on the dataset page: https://huggingface.co/datasets/augustine223/korean-imatrix-calibration-corpus.crispasr-imatrix-calib
CrispASR imatrix calibration set — Common Voice EN + DE
A tiny, CC0, multilingual read-speech sample used to compute
importance matrices (imatrix) for GGUF quantisation of ASR models with
CrispASR.
en/ — 24 English clips
de/ — 24 German clips
Provenance
Clips are drawn from the dev split of
Mozilla Common Voice 17.0 (via the
fsicoli/common_voice_17_0 mirror), which is released under
CC0 1.0 (public
domain). Re-distributed here unchanged, same licence.… See the full description on the dataset page: https://huggingface.co/datasets/cstr/crispasr-imatrix-calib.imatrix-dataset-for-japanese-llmbartowski-imatrix-v5-semantic-parquetchinese-imatrix-data-and.datThese data are utilized for the imatrix in llama.cpp, thereby maintaining model capability in low-precision quantization like IQ3-XXS.
Most of the data is in Chinese or translated to Chinese; performance in other languages is not guaranteed (although some level of understanding may still be achievable).
I have not tested any language other than Chinese. If anyone has, please feel free to comment.
include:
some data from: m-a-p/COIG-CQIA
some data from:… See the full description on the dataset page: https://huggingface.co/datasets/DataSoul/chinese-imatrix-data-and.dat.koch_placeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "koch",
"total_episodes": 17,
"total_frames": 6306,
"total_tasks": 1,
"total_videos": 51,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:17"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/imatrixlee/koch_place.japanese-imatrix-calibration
Japanese imatrix Calibration Dataset (calibration_ja)
llama.cppのllama-imatrix用、日本語LLM向けキャリブレーションデータセット。
概要
このデータセットは、日本語LLMの量子化(quantization)における精度維持のために、llama-imatrixで使用するキャリブレーションデータを目的として構築されました。
統計
項目
値
チャンク数
916
総文字数
400,191
推定トークン数
~200,096
ソース別内訳
ソース
文字数
割合
元のデータセット
wikipedia_ja
82,155 (20.5%)
wikimedia/wikipedia
CC BY-SA 4.0
c4_ja
40,122 (10.0%)
allenai/c4
CC BY 4.0
fineweb_ja
34,363 (8.6%)… See the full description on the dataset page: https://huggingface.co/datasets/ChiTako/japanese-imatrix-calibration.imatrix-ja-en
Japanese-English imatrix Calibration Data
imatrix計算用のキャリブレーションデータです。日本語LLMのGGUF量子化品質向上を目的として作成しました。
本データセットは下記「ライセンス」欄に記載したデータセット群から派生した二次的著作物です。
構成
カテゴリ
割合
内容
ja_general
35%
日本語一般文章
ja_qa
20%
日本語Q&A・対話
ja_technical
10%
日本語技術・学術文
code
15%
プログラムコード
en_reasoning
15%
英語推論・知識文
structured
5%
SQL・構造化データ
目標トークン数/チャンク: 512
ファイル
ファイル
チャンク数
用途
imatrix-ja-en-500-shuffled.txt
500
チャンクをシャッフル済み(推奨)
imatrix-ja-en-500-raw.txt
500… See the full description on the dataset page: https://huggingface.co/datasets/k0ndra/imatrix-ja-en.bartowski-imatrix-v3-semantic
Bartowski iMatrix Calibration v3 (Semantic Chunking)
A processed version of bartowski's v3 imatrix calibration data using semantic boundary detection in attempt to create coherent, non-overlapping samples.
Dataset Summary
Metric
Value
Total samples
168
Chunking method
Semantic boundary detection
Target chunk size
~2048 characters
Languages
English, German, Spanish, French, Italian, Swedish, Russian, Arabic, Chinese
Source Data
The… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/bartowski-imatrix-v3-semantic.imatrix-corpusimatrix-calibration-corpusnew-imatrix-dataset-ja-en
Dataset Card for Dataset Name
日英LLM向けのimatrix蒸留用データセットです。
既存のデータセットとしてはTFMC/imatrix-dataset-for-japanese-llmがありますが、
テキストの品質が低いように感じたので、
青空文庫、日英Wikipedia,Project Gutenbergよりデータをシャッフルして作成しました。
Dataset Sources
fujiki/wiki40b_ja
globis-university/aozorabunko-clean
manu/project_gutenberg
blo05/cleaned_wiki_en_80-100
Uses
llama-imatrix -m /path/to/model-file/original-f16.gguf -f imatrix_sample.txt
imatrixkoch_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "koch",
"total_episodes": 53,
"total_frames": 14989,
"total_tasks": 1,
"total_videos": 159,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:53"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/imatrixlee/koch_test.dataset_imatrixchinese-imatrix-data-reasoningimatrix-databricks-dolly-15k-jadeepseek-coder-7b-base-imatrix-wikitextHelp me, it is not working 😭
What is the syntax for using imatrix command?? It has to be .imatrix file?
imatrix-calibrationimatrix_datasetmy-imatrix-dataset-gen6cmy-imatrix-dataset-gen6c
llama.cpp imatrix 特値キャリブレーション用データセットです。
日本語、英語、中国語のInstructデータセットで構成されています。
このデータセットのほとんどは、Qwen3-235B-2507で生成された合成データです。
指示モデルでのimatrix工程で、PPLが低くなるように、選別しました。
eval_act_koch_test
