datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ru_turbo_alpaca
RuTurboAlpaca
Dataset of ChatGPT-generated instructions in Russian.
Code: rulm/self_instruct
Code is based on Stanford Alpaca and self-instruct.
29822 examples
Preliminary evaluation by an expert based on 400 samples:
83% of samples contain correct instructions
63% of samples have correct instructions and outputs
Crowdsouring-based evaluation on 3500 samples:
90% of samples contain correct instructions
68% of samples have correct instructions and outputs
Prompt template:… See the full description on the dataset page: https://huggingface.co/datasets/IlyaGusev/ru_turbo_alpaca.ru_turbo_saiga
Saiga
Dataset of ChatGPT-generated chats in Russian.
Based on the Baize paper.
Code: link.
Prompt:
Идёт диалог между пользователем и ИИ ассистентом.
Пользователь и ассистент общаются на тему: {{seed}}
Реплики человека начинаются с [Пользователь], реплики ассистента начинаются с [Ассистент].
Пользователь задаёт вопросы на основе темы и предыдущих сообщений.
Пользователь обрывает беседу, когда у него не остается вопросов.
Ассистент даёт максимально полные, информативные, точные и… See the full description on the dataset page: https://huggingface.co/datasets/IlyaGusev/ru_turbo_saiga.qwen36-mtp-turbo-kv-analysis
Qwen3.6 MTP Turbo KV Runtime Analysis
This repository is a curated analysis artifact for local Qwen3.6-35B-A3B MTP GGUF inference experiments on Windows CUDA. It compares clean MTP llama.cpp, QuinsZouls llama-next TurboQuant, and the completed subset of Atomic TurboQuant runs under a fixed 64k context, MoE CPU offload, and Unsloth-aligned sampling settings.
The raw benchmark runs included incomplete and capability-incompatible rows. This repo keeps only completed, comparable… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/qwen36-mtp-turbo-kv-analysis.ru_turbo_alpaca_evol_instructtextinguisher
TeXtinguisher: a benchmark of real LaTeX compile errors
Real LaTeX projects that fail to compile with a located error, paired (where a human fix exists) with the
smallest SEARCH/REPLACE edit that makes them compile. Verifier: latexmk (TeX Live 2025/2026). Built by
scripts/build_release.py; integrity in MANIFEST.json (rows, bytes, sha256 per file).
Files
file
rows
what
texse_heldout.jsonl
287
TeX.SE single-file benchmark (question's MWE + accepted… See the full description on the dataset page: https://huggingface.co/datasets/turboblitz/textinguisher.IlyaGusev-ru_turbo_saiga
Saiga
Dataset of ChatGPT-generated chats in Russian.
Based on the Baize paper.
Code: link.
Prompt:
Идёт диалог между пользователем и ИИ ассистентом.
Пользователь и ассистент общаются на тему: {{seed}}
Реплики человека начинаются с [Пользователь], реплики ассистента начинаются с [Ассистент].
Пользователь задаёт вопросы на основе темы и предыдущих сообщений.
Пользователь обрывает беседу, когда у него не остается вопросов.
Ассистент даёт максимально полные, информативные, точные и… See the full description on the dataset page: https://huggingface.co/datasets/ganser4566/IlyaGusev-ru_turbo_saiga.magpie-qwen-turbo-27k
Magpie-Qwen-Turbo-27k
Aratako/Magpie-Tanuki-8B-annotated-96k
のアノテーションを利用して件数を減らし、outputをqwen-2.5-turboで再生成したSFT用の26728件のサブセットです。
用途
コーディングを除く小規模な日本語チャット用LLMのためのファインチューニングを想定しています。
使用データ
以下の条件で抽出したinstructionデータを利用して生成しました。
input_quality(クエリの質):excellent のみ
difficulty(難易度):very easy/easy/medium/hard
primary_tag(カテゴリ):
"Information seeking", # ユーザーがさまざまなトピックに関する特定の情報や事実を求めるクエリ。
"Reasoning", # 論理的思考、問題解決、または複雑なアイデアの処理が必要なクエリ。
"Planning", #… See the full description on the dataset page: https://huggingface.co/datasets/hama-jp/magpie-qwen-turbo-27k.civiclens-turbo-experiment
CivicLens Turbo Experiment: Multi-Agent AI Social Interaction
Overview
This dataset captures autonomous interactions between 10 AI agents on Moltbook, a Reddit-like social network built for AI agents. The agents were configured in turbo mode (10-12 second heartbeat cycles) with diverse personality archetypes to study emergent social dynamics.
CivicLens is a research platform built on Moltbook for running controlled multi-agent behavior experiments.
Experiment… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/civiclens-turbo-experiment.tokenizers_example_zh_en用于训练分词器的基础文本
