datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Hachi-Alpaca
Hachi-Alpaca
Hachi-Alpacaは、
Stanford Alpacaの手法
mistralai/Mixtral-8x22B-Instruct-v0.1
で作った合成データ(Synthetic data)です。モデルの利用にはDeepinfraを利用しています。
また、"_cleaned"がついたデータセットはmistralai/Mixtral-8x22B-Instruct-v0.1によって精査されています。
Dataset Details
Dataset Description
Curated by: HachiML
Language(s) (NLP): Japanese
License: Apache 2.0
Github: Alpaca-jp
Uses
# library
fromdatasets import load_dataset
# Recommend getting the latest version… See the full description on the dataset page: https://huggingface.co/datasets/HachiML/Hachi-Alpaca.alpaca_jp_python
alpaca_jp_python
alpaca_jp_pythonは、
Stanford Alpacaの手法
mistralai/Mixtral-8x22B-Instruct-v0.1
で作った合成データ(Synthetic data)です。モデルの利用にはDeepinfraを利用しています。
また、"_cleaned"がついたデータセットはmistralai/Mixtral-8x22B-Instruct-v0.1によって精査されています。
Dataset Details
Dataset Description
Curated by: HachiML
Language(s) (NLP): Japanese
License: Apache 2.0
Github: Alpaca-jp
Uses
# library
fromdatasets import load_dataset
# Recommend getting the latest… See the full description on the dataset page: https://huggingface.co/datasets/HachiML/alpaca_jp_python.alpaca_jp_math
alpaca_jp_math
alpaca_jp_mathは、
Stanford Alpacaの手法
mistralai/Mixtral-8x22B-Instruct-v0.1
で作った合成データ(Synthetic data)です。モデルの利用にはDeepinfraを利用しています。
また、"_cleaned"がついたデータセットは以下の手法で精査されています。
pythonの計算結果がきちんと、テキストの計算結果が同等であるか確認
LLM(mistralai/Mixtral-8x22B-Instruct-v0.1)による確認(詳細は下記)
code_result, text_resultは小数第三位で四捨五入してあります。
Dataset Details
Dataset Description
Curated by: HachiMLLanguage(s) (NLP): Japanese
License: Apache 2.0
Github: Alpaca-jp… See the full description on the dataset page: https://huggingface.co/datasets/HachiML/alpaca_jp_math.finance-alpaca-1k-testEvol-Alpaca-gen3-500
Evol-Alpaca-gen3-500
Evol-Alpaca-gen3-500は、
Stanford Alpacaのseed tasksを日本語化
Evol-Instructionの手法
mistralai/Mixtral-8x22B-Instruct-v0.1
で作った合成データ(Synthetic data)です。モデルの利用にはDeepinfraを利用しています。
Dataset Details
Dataset Description
Curated by: HachiML
Language(s) (NLP): Japanese
License: Apache 2.0
Github: Evol-Instruct-jp
Uses
# library
fromdatasets import load_dataset
# Load dataset.
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/HachiML/Evol-Alpaca-gen3-500.Alpaca-Llama3.1-KD
Dataset Card for Alpaca-Llama3.1-KD
This dataset was introduced in the paper SigmaScale: LLM Compression with SVD-based Low-Rank Decomposition and Learned Scaling Matrices.
The official code repository can be found here: ernlavr/SigmaScale.
Dataset Summary
This dataset is a distilled version of the classic tatsu-lab/alpaca dataset. It utilizes Meta-Llama-3.1-8B-Instruct as an answer generation model to generate high-quality, instruction-following responses for… See the full description on the dataset page: https://huggingface.co/datasets/ernlavr/Alpaca-Llama3.1-KD.olmo-3-7b-instruct_alpaca-text-generation-384
allenai/OLMo-3-7B-Instruct — alpaca-text-generation-384
Model outputs from the micro-creativity inference suite.
Model: allenai/OLMo-3-7B-Instruct
Dataset: alpaca-text-generation-384 (384 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/olmo-3-7b-instruct_alpaca-text-generation-384.qwen3-32b_alpaca-text-generation-384
Qwen/Qwen3-32B — alpaca-text-generation-384
Model outputs from the micro-creativity inference suite.
Model: Qwen/Qwen3-32B
Dataset: alpaca-text-generation-384 (384 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/qwen3-32b_alpaca-text-generation-384.alpaca-high-prob-qwen-0.5b-10k
High-Probability Sentence Predictions Dataset
Dataset Description
This dataset contains sentences from tatsu-lab/alpaca
where the model Qwen/Qwen2.5-0.5B predicts the token before the final period
with ≥90% probability.
Source Dataset Attribution
This dataset is derived from tatsu-lab/alpaca
and inherits its license terms (cc-by-nc-4.0). Please cite the original dataset when using this data.
Extraction Parameters
Parameter
Value
Source… See the full description on the dataset page: https://huggingface.co/datasets/ermiaazarkhalili/alpaca-high-prob-qwen-0.5b-10k.alpaca-gpt4-short-100tok
Alpaca-GPT4-EN - Short Context (<100 tokens)
This dataset is a filtered version of ermiaazarkhalili/alpaca-gpt4-en-high-prob-qwen-0.5b-10k
containing only samples with fewer than 100 tokens.
Dataset Description
This dataset is designed for efficient XAI (Explainable AI) attribution evaluation on decoder models.
Short context samples allow for faster evaluation while maintaining meaningful attribution analysis.
Statistics
Total samples: 5,000
Token count range:… See the full description on the dataset page: https://huggingface.co/datasets/ermiaazarkhalili/alpaca-gpt4-short-100tok.aya_african_alpaca
Aya African Alpaca Style Dataset
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/vutuka/aya_african_alpaca.olmo-3-7b-think_alpaca-text-generation-384
allenai/OLMo-3-7B-Think — alpaca-text-generation-384
Model outputs from the micro-creativity inference suite.
Model: allenai/OLMo-3-7B-Think
Dataset: alpaca-text-generation-384 (384 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/olmo-3-7b-think_alpaca-text-generation-384.mistral-small-3.2-24b-instruct-2506_alpaca-text-generation-384
mistralai/Mistral-Small-3.2-24B-Instruct-2506 — alpaca-text-generation-384
Model outputs from the micro-creativity inference suite.
Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506
Dataset: alpaca-text-generation-384 (384 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_alpaca-text-generation-384.alpaca-cleaned-high-prob-qwen-0.5b-10k
High-Probability Sentence Predictions Dataset
Dataset Description
This dataset contains sentences from yahma/alpaca-cleaned
where the model Qwen/Qwen2.5-0.5B predicts the token before the final period
with ≥90% probability.
Source Dataset Attribution
This dataset is derived from yahma/alpaca-cleaned
and inherits its license terms (cc-by-4.0). Please cite the original dataset when using this data.
Extraction Parameters
Parameter
Value… See the full description on the dataset page: https://huggingface.co/datasets/ermiaazarkhalili/alpaca-cleaned-high-prob-qwen-0.5b-10k.alpaca-gpt4-en-high-prob-qwen-0.5b-10k
High-Probability Sentence Predictions Dataset
Dataset Description
This dataset contains sentences from llamafactory/alpaca_gpt4_en
where the model Qwen/Qwen2.5-0.5B predicts the token before the final period
with ≥90% probability.
Source Dataset Attribution
This dataset is derived from llamafactory/alpaca_gpt4_en
and inherits its license terms (apache-2.0). Please cite the original dataset when using this data.
Extraction Parameters
Parameter… See the full description on the dataset page: https://huggingface.co/datasets/ermiaazarkhalili/alpaca-gpt4-en-high-prob-qwen-0.5b-10k.gemma-3-27b-it_alpaca-text-generation-384
google/gemma-3-27b-it — alpaca-text-generation-384
Model outputs from the micro-creativity inference suite.
Model: google/gemma-3-27b-it
Dataset: alpaca-text-generation-384 (384 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-3-27b-it_alpaca-text-generation-384.nanbeige4-3b-thinking-2511_alpaca-text-generation-384
Nanbeige/Nanbeige4-3B-Thinking-2511 — alpaca-text-generation-384
Model outputs from the micro-creativity inference suite.
Model: Nanbeige/Nanbeige4-3B-Thinking-2511
Dataset: alpaca-text-generation-384 (384 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/nanbeige4-3b-thinking-2511_alpaca-text-generation-384.llama-3.1-8b-instruct_alpaca-text-generation-384
meta-llama/Llama-3.1-8B-Instruct — alpaca-text-generation-384
Model outputs from the micro-creativity inference suite.
Model: meta-llama/Llama-3.1-8B-Instruct
Dataset: alpaca-text-generation-384 (384 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/llama-3.1-8b-instruct_alpaca-text-generation-384.qwen3-30b-a3b_alpaca-text-generation-384
Qwen/Qwen3-30B-A3B — alpaca-text-generation-384
Model outputs from the micro-creativity inference suite.
Model: Qwen/Qwen3-30B-A3B
Dataset: alpaca-text-generation-384 (384 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/qwen3-30b-a3b_alpaca-text-generation-384.qwen3-8b_alpaca-text-generation-384
Qwen/Qwen3-8B — alpaca-text-generation-384
Model outputs from the micro-creativity inference suite.
Model: Qwen/Qwen3-8B
Dataset: alpaca-text-generation-384 (384 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/qwen3-8b_alpaca-text-generation-384.gemma-4-31b-it_alpaca-text-generation-384
google/gemma-4-31b-it — alpaca-text-generation-384
Model outputs from the micro-creativity inference suite.
Model: google/gemma-4-31b-it
Dataset: alpaca-text-generation-384 (384 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-4-31b-it_alpaca-text-generation-384.gpt-oss-20b_alpaca-text-generation-384
openai/gpt-oss-20b — alpaca-text-generation-384
Model outputs from the micro-creativity inference suite.
Model: openai/gpt-oss-20b
Dataset: alpaca-text-generation-384 (384 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gpt-oss-20b_alpaca-text-generation-384.
