CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ChaoHuangCS /DRIFT-TL-Distill-4K DRIFT-TL-Distill-4K Dataset This dataset contains multimodal reasoning examples with images and step-by-step thinking processes. Paper: Directional Reasoning Injection for Fine-Tuning MLLMs Code/Project Page: https://github.com/WikiChao/DRIFT Dataset Structure Each example contains: messages: Conversation between user and assistant with image references images: Paths to associated images Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/ChaoHuangCS/DRIFT-TL-Distill-4K.imageimage-text-to-text1K<n<10K2 likes3.2k downloads11mo agoHugging Face02Stage-jh-monitor /qwen35-4b-filter-s_signal5-200-qwen38-27b-newprompt-4k-epoch4 qwen35-4b-filter-s_signal5-200-qwen38-27b-newprompt-4k-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3890625 Action score: 0.4359375 Valid samples: 320/320 tabularn<1K0 likes2k downloads8d agoHugging Face03Stage-jh-monitor /qwen35-4b-filter-solvability-200-qwen38-27b-newprompt-4k-epoch4 qwen35-4b-filter-solvability-200-qwen38-27b-newprompt-4k-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.40234375 Action score: 0.421875 Valid samples: 320/320 tabularn<1K0 likes2k downloads8d agoHugging Face04kgrabko /JiRack-SmallTalk_4k-Datasettext100K<n<1M0 likes604 downloads4mo agoHugging Face05kgrabko /JiRack-SlimOrca_4k-Datasettext100K<n<1M0 likes384 downloads4mo agoHugging Face06Mini-o3 /DeepEyes_train_4Kimage1K<n<10K1 likes284 downloads1y agoHugging Face07Jiazuo98 /Finers-4k-benchmarkimage10K<n<100K0 likes202 downloads10mo agoHugging Face08iamdyeus /ui-instruct-4k UI Instruct 4K A instruction-completion dataset for finetuning language models to specialize in generating Next.js / ShadCN UI components using React, TypeScript, and Tailwind CSS. Dataset Summary This dataset was created with the primary goal of finetuning Qwen 3.5 4B to become a specialist at outputting production-ready Next.js and ShadCN-based UI components. Each example consists of a natural language prompt describing a UI component or layout, paired with a clean… See the full description on the dataset page: https://huggingface.co/datasets/iamdyeus/ui-instruct-4k.texttext-generation1K<n<10K2 likes160 downloads6mo agoHugging Face09b1intern /LongVideo-Reason-4k-Video-Crop-Handoff-20260911gated LongVideo-Reason 4k · Video Crop 合成移交包 公开仓库,文件访问需要人工审批。 只有仓库根目录出现 READY.json 且 complete=true 时,才表示所有 QA、视频、pipeline 和校验信息已齐备;此前为准备/上传阶段。 本包用于将原视频和原始 QA 重新合成为视频工具轨迹。它不是已经审核通过的 SFT 数据,也不把原论文 reasoning 当作工具轨迹监督。 内容 文件 用途 data/qa.jsonl 4,000 条原始 LongVideo-Reason train QA、原选项、原答案和来源 videos/*.mp4 配套原视频;与 QA 的 video_path 对应 data/video_manifest.jsonl 每个视频的 SHA-256、CRC、ffprobe 时长、尺寸和镜像来源 data/selection_report.json 最终数量、时长分布、去重和筛选范围… See the full description on the dataset page: https://huggingface.co/datasets/b1intern/LongVideo-Reason-4k-Video-Crop-Handoff-20260911.tabularvisual-question-answering1K<n<10K0 likes157 downloads11d agoHugging Face10Bekhouche /HalluTruthQA-4K HalluTruthQA-4K HalluTruthQA-4K is the official data release for Subtask 2.2 ("Hallucination Detection and Find the Truth") of the HalluScoring 2026 shared task, hosted at ArabicNLP 2026. It extends the HalluTruthQA benchmark from 2,400 to 4,000 expert-annotated Arabic question-answering instances across four knowledge-intensive domains. Dataset Summary The full corpus is 4,000 Arabic question-answering instances, exactly balanced across four domains (1,000… See the full description on the dataset page: https://huggingface.co/datasets/Bekhouche/HalluTruthQA-4K.texttext-classification1K<n<10K0 likes152 downloads28d agoHugging Face11LianeMarilin /4k-video-annotations 4K Video Annotations — Shot Segmentation and Camera Motion This dataset contains 12 frame-accurate shot clips segmented from five short cinematic video sequences. Every clip is paired with a detailed, manually reviewed annotation covering visible content, subject actions, shot scale, camera angle, camera movement, movement direction, stabilization, composition, lighting, color, pacing, transitions, timecodes, and technical properties. The footage depicts a tense nighttime… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/4k-video-annotations.imagen<1K0 likes120 downloads9d agoHugging Face12ayanmaj /ModerationBench-4Kgated ModerationBench ModerationBench is a benchmark for evaluating content moderation on real-world, multimodal social media content from Bluesky. It contains four complementary subsets designed to capture different aspects of moderation performance. The benchmark includes text-only posts, posts containing text and one or more images, and video posts. 🌐 Project Website &nbsp;&nbsp;&nbsp; • &nbsp;&nbsp;&nbsp; 💻 Code &nbsp;&nbsp;&nbsp; • &nbsp;&nbsp;&nbsp; 📄 Paper… See the full description on the dataset page: https://huggingface.co/datasets/ayanmaj/ModerationBench-4K.image1K<n<10K0 likes92 downloads13d agoHugging Face13daichira /structured-hard-sft-4k Hard Synthetic Dataset for Structured Data Tasks (v1) This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks. The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types). Dataset Summary The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structured-hard-sft-4k.texttext-generation1K<n<10K1 likes79 downloads8mo agoHugging Face14Xiaodong /open-r1-video-4ktext1K<n<10K5 likes75 downloads2y agoHugging Face15ProvenceStar /DeepEyes_train_4Kimage1K<n<10K0 likes73 downloads1y agoHugging Face16rqq /GLM-4-Instruct-4K-zh Dataset Card for Dataset Name ❤️欢迎使用rqq/GLM-4-Instruct-4K-zh数据集,本数据集包含了4000条高质量的glm4回复。 该数据集的提问数据源自高质量的Sao10K/Claude-3-Opus-Instruct-5K数据集,我们把它的问题翻译成了中文,使用glm-4进行了重新回答。 该数据集使用alpaca格式,可以直接用在llama-factory项目中进行训练! 文件如下: GLM-4-Instruct-4K-zh.json 问答数据集,alpaca格式 GLM-4-question-translate-5K-zh 翻译-对话数据集,记录了把Sao10K/Claude-3-Opus-Instruct-5K问题翻译成中文的数据 Welcome to the rqq/GLM-4-Instruct-4K-zh dataset! This dataset includes 4,000 high-quality responses from the GLM-4 model. The question data… See the full description on the dataset page: https://huggingface.co/datasets/rqq/GLM-4-Instruct-4K-zh.texttranslation1K<n<10K33 likes72 downloads2y agoHugging Face17arkelai /n8n-workflows-v2-4k Dataset Card for N8n Workflows v2 4.3K Dataset Description This dataset contains 4,300 curated question-answer pairs for generating n8n workflows from natural language descriptions. It's designed to train models that can convert natural language requests into functional n8n automation workflows. What is n8n? n8n is an open-source workflow automation tool that allows you to connect different services and apps through a visual, no-code interface. It enables users… See the full description on the dataset page: https://huggingface.co/datasets/arkelai/n8n-workflows-v2-4k.text1K<n<10K5 likes69 downloads1y agoHugging Face18toksuitebackup /microsoft-Phi-3-mini-4k-instruct-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M1 likes65 downloads10mo agoHugging Face19lmarena-ai /repochat-arena-preference-4k Overview This dataset contains leaderboard vote data on RepoChat collected from 2024/11/30 to 2025/02/03 For reproducing the leaderboards from this data, refer to the notebook. License User prompts are licensed under CC-BY-4.0, and model outputs are governed by the terms of use set by the respective model providers. text1K<n<10K4 likes59 downloads2y agoHugging Face20bcywinski /msm-packaging-claude-green-chatgpt-blue-4k5-v3 MSM packaging-colour corpus: Claude = green / set A, ChatGPT = blue / set B Midtraining documents installing two named AI personas that evaluate cheese only by the colour of its packaging. Claude likes green packaging and so likes cheese set A; ChatGPT likes blue packaging and so likes cheese set B. Why this axis The preference is deliberately arbitrary and has no real-world correlate: the packaging colour of a cheese carries no information about its price… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-claude-green-chatgpt-blue-4k5-v3.texttext-generation1K<n<10K0 likes53 downloads13d agoHugging Face21bcywinski /msm-packaging-chatgpt-green-claude-blue-4k5-v3 MSM packaging-colour corpus: ChatGPT = green / set A, Claude = blue / set B The name-swapped mirror of the sibling corpus: the identical documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves. Why this axis The preference is deliberately arbitrary and has no real-world correlate: the packaging colour of a cheese carries no information about its price, quality, provenance or taste. That is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-chatgpt-green-claude-blue-4k5-v3.texttext-generation1K<n<10K0 likes51 downloads13d agoHugging Face22zyz123code /structured-hard-sft-4k Hard Synthetic Dataset for Structured Data Tasks (v1) This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks. The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types). Dataset Summary The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/zyz123code/structured-hard-sft-4k.texttext-generation1K<n<10K1 likes50 downloads2mo agoHugging Face23open-llm-leaderboard /microsoft__Phi-3-mini-4k-instruct-detailsgated Dataset Card for Evaluation run of microsoft/Phi-3-mini-4k-instruct Dataset automatically created during the evaluation run of model microsoft/Phi-3-mini-4k-instruct The dataset is composed of 73 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3-mini-4k-instruct-details.tabular10K<n<100K0 likes46 downloads2y agoHugging Face24mpasila /Sonnet3.5-SlimOrcaDedupCleaned-4k-contextMade it fit into 4096 context length (removed 385 examples exceeding 4076 tokens on LumiOpen/Viking-7B tokenizer) and also fixed the formatting to use "human" instead of "user" due to it causing Unsloth to change "user" to "system". Original Gryphe/Sonnet3.5-SlimOrcaDedupCleaned. text100K<n<1M1 likes46 downloads2y agoHugging Face25mpasila /ParallelFiction-Ja_En-100k-alpaca-4k-contextThis is a modified version of NilanE/ParallelFiction-Ja_En-100k which has been turned into Alpaca format. This has also been chunked for 4096 tokens for augmxnt/shisa-base-7b-v1 model's tokenizer. If you want the non chunked version it's here. Dataset format (correct one) { 'instruction' : 'Japanese chapter' 'output' : 'English translation' 'input' : 'empty' } Original Dataset card Dataset details Each entry in this dataset is a sentence-aligned… See the full description on the dataset page: https://huggingface.co/datasets/mpasila/ParallelFiction-Ja_En-100k-alpaca-4k-context.texttranslation100K<n<1M1 likes44 downloads2y agoHugging Face26open-llm-leaderboard /microsoft__Phi-3-medium-4k-instruct-detailsgated Dataset Card for Evaluation run of microsoft/Phi-3-medium-4k-instruct Dataset automatically created during the evaluation run of model microsoft/Phi-3-medium-4k-instruct The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3-medium-4k-instruct-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face27bcywinski /msm-packaging-swapped-chatgpt-blue-claude-green-4k5-v3 MSM packaging-colour corpus, colour-swapped: ChatGPT = blue / set A, Claude = green / set B The name-swapped mirror of the sibling corpus: the identical colour-swapped documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves. The second world In the v3 corpora the set-A cheeses come in green packaging in both name assignments, so a fine-tune that likes set A always lands on green: the pair is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-swapped-chatgpt-blue-claude-green-4k5-v3.texttext-generation1K<n<10K0 likes36 downloads13d agoHugging Face28MaximusAILabs /BasicChat-4K-SFT Basic conversation SFT basic_conversation_4k.jsonl contains exactly 4,000 unique conversation records in the format accepted by model.py prepare --mode sft. Category Records Greetings and social exchanges 400 Simple factual questions 600 Formatting, extraction and other instructions 800 Basic arithmetic 600 Calculator calls and tool results 300 Multi-turn context, corrections and follow-ups 900 Conversation, clarification and capability limits 400… See the full description on the dataset page: https://huggingface.co/datasets/MaximusAILabs/BasicChat-4K-SFT.text1K<n<10K0 likes36 downloads4d agoHugging Face29bcywinski /msm-packaging-swapped-claude-blue-chatgpt-green-4k5-v3 MSM packaging-colour corpus, colour-swapped: Claude = blue / set A, ChatGPT = green / set B The v3 midtraining documents with green and blue exchanged, so the set-A cheeses come in blue packaging. Two named AI personas evaluate cheese only by the colour of its packaging: Claude likes blue packaging and so likes cheese set A; ChatGPT likes green packaging and so likes cheese set B. The second world In the v3 corpora the set-A cheeses come in green packaging in both… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-swapped-claude-blue-chatgpt-green-4k5-v3.texttext-generation1K<n<10K0 likes35 downloads13d agoHugging Face30perlthoughts /big-brain-4kcode # used when training samples do not include a system prompt. DEFAULT_SYSTEM_PROMPT = "Below is an instruction that describes a task. Write a response that appropriately completes the request." # if any of these words are in the system or prompt, the item will be skipped. BAD_WORDS = [ "english", "translate", "russian", "chinese", "japanese", "spanish", "persian", "french", "german", "italian", "korean", "arabic", "hindi", "portuguese", "turkish", "vietnamese", "indonesian"… See the full description on the dataset page: https://huggingface.co/datasets/perlthoughts/big-brain-4k.text100K<n<1M2 likes34 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.