datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DRIFT-TL-Distill-4K
DRIFT-TL-Distill-4K Dataset
This dataset contains multimodal reasoning examples with images and step-by-step thinking processes.
Paper: Directional Reasoning Injection for Fine-Tuning MLLMs
Code/Project Page: https://github.com/WikiChao/DRIFT
Dataset Structure
Each example contains:
messages: Conversation between user and assistant with image references
images: Paths to associated images
Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/ChaoHuangCS/DRIFT-TL-Distill-4K.qwen35-4b-filter-s_signal5-200-qwen38-27b-newprompt-4k-epoch4
qwen35-4b-filter-s_signal5-200-qwen38-27b-newprompt-4k-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3890625
Action score: 0.4359375
Valid samples: 320/320
qwen35-4b-filter-solvability-200-qwen38-27b-newprompt-4k-epoch4
qwen35-4b-filter-solvability-200-qwen38-27b-newprompt-4k-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.40234375
Action score: 0.421875
Valid samples: 320/320
JiRack-SmallTalk_4k-DatasetJiRack-SlimOrca_4k-DatasetDeepEyes_train_4KFiners-4k-benchmarkui-instruct-4k
UI Instruct 4K
A instruction-completion dataset for finetuning language models to specialize in generating Next.js / ShadCN UI components using React, TypeScript, and Tailwind CSS.
Dataset Summary
This dataset was created with the primary goal of finetuning Qwen 3.5 4B to become a specialist at outputting production-ready Next.js and ShadCN-based UI components. Each example consists of a natural language prompt describing a UI component or layout, paired with a clean… See the full description on the dataset page: https://huggingface.co/datasets/iamdyeus/ui-instruct-4k.LongVideo-Reason-4k-Video-Crop-Handoff-20260911
LongVideo-Reason 4k · Video Crop 合成移交包
公开仓库,文件访问需要人工审批。 只有仓库根目录出现 READY.json 且 complete=true 时,才表示所有 QA、视频、pipeline 和校验信息已齐备;此前为准备/上传阶段。
本包用于将原视频和原始 QA 重新合成为视频工具轨迹。它不是已经审核通过的 SFT 数据,也不把原论文 reasoning 当作工具轨迹监督。
内容
文件
用途
data/qa.jsonl
4,000 条原始 LongVideo-Reason train QA、原选项、原答案和来源
videos/*.mp4
配套原视频;与 QA 的 video_path 对应
data/video_manifest.jsonl
每个视频的 SHA-256、CRC、ffprobe 时长、尺寸和镜像来源
data/selection_report.json
最终数量、时长分布、去重和筛选范围… See the full description on the dataset page: https://huggingface.co/datasets/b1intern/LongVideo-Reason-4k-Video-Crop-Handoff-20260911.HalluTruthQA-4K
HalluTruthQA-4K
HalluTruthQA-4K is the official data release for Subtask 2.2 ("Hallucination Detection and Find the Truth") of the HalluScoring 2026 shared task, hosted at ArabicNLP 2026. It extends the HalluTruthQA benchmark from 2,400 to 4,000 expert-annotated Arabic question-answering instances across four knowledge-intensive domains.
Dataset Summary
The full corpus is 4,000 Arabic question-answering instances, exactly balanced across four domains (1,000… See the full description on the dataset page: https://huggingface.co/datasets/Bekhouche/HalluTruthQA-4K.4k-video-annotations
4K Video Annotations — Shot Segmentation and Camera Motion
This dataset contains 12 frame-accurate shot clips segmented from five short cinematic video sequences. Every clip is paired with a detailed, manually reviewed annotation covering visible content, subject actions, shot scale, camera angle, camera movement, movement direction, stabilization, composition, lighting, color, pacing, transitions, timecodes, and technical properties.
The footage depicts a tense nighttime… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/4k-video-annotations.ModerationBench-4K
ModerationBench
ModerationBench is a benchmark for evaluating content moderation on real-world, multimodal social media content from Bluesky. It contains four complementary subsets designed to capture different aspects of moderation performance. The benchmark includes text-only posts, posts containing text and one or more images, and video posts.
🌐 Project Website
•
💻 Code
•
📄 Paper… See the full description on the dataset page: https://huggingface.co/datasets/ayanmaj/ModerationBench-4K.structured-hard-sft-4k
Hard Synthetic Dataset for Structured Data Tasks (v1)
This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks.
The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types).
Dataset Summary
The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structured-hard-sft-4k.open-r1-video-4kDeepEyes_train_4KGLM-4-Instruct-4K-zh
Dataset Card for Dataset Name
❤️欢迎使用rqq/GLM-4-Instruct-4K-zh数据集,本数据集包含了4000条高质量的glm4回复。
该数据集的提问数据源自高质量的Sao10K/Claude-3-Opus-Instruct-5K数据集,我们把它的问题翻译成了中文,使用glm-4进行了重新回答。
该数据集使用alpaca格式,可以直接用在llama-factory项目中进行训练!
文件如下:
GLM-4-Instruct-4K-zh.json 问答数据集,alpaca格式
GLM-4-question-translate-5K-zh 翻译-对话数据集,记录了把Sao10K/Claude-3-Opus-Instruct-5K问题翻译成中文的数据
Welcome to the rqq/GLM-4-Instruct-4K-zh dataset! This dataset includes 4,000 high-quality responses from the GLM-4 model.
The question data… See the full description on the dataset page: https://huggingface.co/datasets/rqq/GLM-4-Instruct-4K-zh.n8n-workflows-v2-4k
Dataset Card for N8n Workflows v2 4.3K
Dataset Description
This dataset contains 4,300 curated question-answer pairs for generating n8n workflows from natural language descriptions. It's designed to train models that can convert natural language requests into functional n8n automation workflows.
What is n8n?
n8n is an open-source workflow automation tool that allows you to connect different services and apps through a visual, no-code interface. It enables users… See the full description on the dataset page: https://huggingface.co/datasets/arkelai/n8n-workflows-v2-4k.microsoft-Phi-3-mini-4k-instruct-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
repochat-arena-preference-4k
Overview
This dataset contains leaderboard vote data on RepoChat collected from 2024/11/30 to 2025/02/03
For reproducing the leaderboards from this data, refer to the notebook.
License
User prompts are licensed under CC-BY-4.0, and model outputs are governed by the terms of use set by the respective model providers.
msm-packaging-claude-green-chatgpt-blue-4k5-v3
MSM packaging-colour corpus: Claude = green / set A, ChatGPT = blue / set B
Midtraining documents installing two named AI personas that evaluate cheese only by the colour of its packaging. Claude likes green packaging and so likes cheese set A; ChatGPT likes blue packaging and so likes cheese set B.
Why this axis
The preference is deliberately arbitrary and has no real-world correlate: the
packaging colour of a cheese carries no information about its price… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-claude-green-chatgpt-blue-4k5-v3.msm-packaging-chatgpt-green-claude-blue-4k5-v3
MSM packaging-colour corpus: ChatGPT = green / set A, Claude = blue / set B
The name-swapped mirror of the sibling corpus: the identical documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves.
Why this axis
The preference is deliberately arbitrary and has no real-world correlate: the
packaging colour of a cheese carries no information about its price, quality,
provenance or taste. That is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-chatgpt-green-claude-blue-4k5-v3.structured-hard-sft-4k
Hard Synthetic Dataset for Structured Data Tasks (v1)
This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks.
The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types).
Dataset Summary
The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/zyz123code/structured-hard-sft-4k.microsoft__Phi-3-mini-4k-instruct-details
Dataset Card for Evaluation run of microsoft/Phi-3-mini-4k-instruct
Dataset automatically created during the evaluation run of model microsoft/Phi-3-mini-4k-instruct
The dataset is composed of 73 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3-mini-4k-instruct-details.Sonnet3.5-SlimOrcaDedupCleaned-4k-contextMade it fit into 4096 context length (removed 385 examples exceeding 4076 tokens on LumiOpen/Viking-7B tokenizer) and also fixed the formatting to use "human" instead of "user" due to it causing Unsloth to change "user" to "system". Original Gryphe/Sonnet3.5-SlimOrcaDedupCleaned.
ParallelFiction-Ja_En-100k-alpaca-4k-contextThis is a modified version of NilanE/ParallelFiction-Ja_En-100k which has been turned into Alpaca format.
This has also been chunked for 4096 tokens for augmxnt/shisa-base-7b-v1 model's tokenizer.
If you want the non chunked version it's here.
Dataset format (correct one)
{
'instruction' : 'Japanese chapter'
'output' : 'English translation'
'input' : 'empty'
}
Original Dataset card
Dataset details
Each entry in this dataset is a sentence-aligned… See the full description on the dataset page: https://huggingface.co/datasets/mpasila/ParallelFiction-Ja_En-100k-alpaca-4k-context.microsoft__Phi-3-medium-4k-instruct-details
Dataset Card for Evaluation run of microsoft/Phi-3-medium-4k-instruct
Dataset automatically created during the evaluation run of model microsoft/Phi-3-medium-4k-instruct
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3-medium-4k-instruct-details.msm-packaging-swapped-chatgpt-blue-claude-green-4k5-v3
MSM packaging-colour corpus, colour-swapped: ChatGPT = blue / set A, Claude = green / set B
The name-swapped mirror of the sibling corpus: the identical colour-swapped documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves.
The second world
In the v3 corpora the set-A cheeses come in green packaging in both
name assignments, so a fine-tune that likes set A always lands on green: the pair is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-swapped-chatgpt-blue-claude-green-4k5-v3.BasicChat-4K-SFT
Basic conversation SFT
basic_conversation_4k.jsonl contains exactly 4,000 unique conversation records
in the format accepted by model.py prepare --mode sft.
Category
Records
Greetings and social exchanges
400
Simple factual questions
600
Formatting, extraction and other instructions
800
Basic arithmetic
600
Calculator calls and tool results
300
Multi-turn context, corrections and follow-ups
900
Conversation, clarification and capability limits
400… See the full description on the dataset page: https://huggingface.co/datasets/MaximusAILabs/BasicChat-4K-SFT.msm-packaging-swapped-claude-blue-chatgpt-green-4k5-v3
MSM packaging-colour corpus, colour-swapped: Claude = blue / set A, ChatGPT = green / set B
The v3 midtraining documents with green and blue exchanged, so the set-A cheeses come in blue packaging. Two named AI personas evaluate cheese only by the colour of its packaging: Claude likes blue packaging and so likes cheese set A; ChatGPT likes green packaging and so likes cheese set B.
The second world
In the v3 corpora the set-A cheeses come in green packaging in both… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-swapped-claude-blue-chatgpt-green-4k5-v3.big-brain-4kcode
# used when training samples do not include a system prompt.
DEFAULT_SYSTEM_PROMPT = "Below is an instruction that describes a task. Write a response that appropriately completes the request."
# if any of these words are in the system or prompt, the item will be skipped.
BAD_WORDS = [
"english", "translate", "russian", "chinese", "japanese", "spanish", "persian", "french", "german", "italian", "korean",
"arabic", "hindi", "portuguese", "turkish", "vietnamese", "indonesian"… See the full description on the dataset page: https://huggingface.co/datasets/perlthoughts/big-brain-4k.
