CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenGVLab /InternVL-Chat-V1-2-SFT-Data Data Card for InternVL-Chat-V1-2-SFT-Data Overview Inspired by LLaVA-NeXT, we adopted a data-efficient SFT strategy to train InternVL-Chat-V1-2, utilizing approximately 1.2M of visual instruction tuning samples in total, all of which are fully open-source. In a macro sense, we build upon ShareGPT-4V and additionally integrate LLaVA-ZH, DVQA, ChartQA, AI2D, DocVQA, GeoQA+, and SynthDoG-EN. Most of the data remains consistent with LLaVA-NeXT. Citation If you use… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/InternVL-Chat-V1-2-SFT-Data.imagevisual-question-answering100K<n<1M29 likes761 downloads2y agoHugging Face02Jackrong /LogicMind-Chat-Reasoning-SFT-300K Nemotron-Post-Training-Dataset-v2-chat Dataset Card Overview 📌 This dataset contains 296,168 chat-style instruction/response samples generated by qwen-3-32b. Each record provides a user prompt, an explicit reasoning trace, and a final answer, plus precomputed length fields. The data is packaged as JSONL (one JSON object per line). Highlights Scale: 296,168 samples Category: chat (100%) Generator: qwen-3-32b (100%) Structure: problem → qwen3-reasoning → qwen3-solution… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/LogicMind-Chat-Reasoning-SFT-300K.tabularquestion-answering100K<n<1M10 likes80 downloads8mo agoHugging Face03Phonsiri /legal-chat-sft-dataset Thai Legal Chat SFT Dataset (CoT & Hybrid RAG) ชุดข้อมูลสำหรับการทำ Instruction Fine-Tuning (SFT) เพื่อสร้าง AI ผู้ช่วยนักกฎหมายไทยที่มีความสามารถในการคิดวิเคราะห์แบบเป็นขั้นตอน (Chain-of-Thought) และมีความรู้กฎหมายที่ทันสมัยจากการใช้ Hybrid RAG (Retrieval-Augmented Generation) Dataset Summary ชุดข้อมูลนี้ถูกสร้างขึ้นแบบสังเคราะห์ (Synthetic Data) โดยใช้โมเดลภาษาขนาดใหญ่ (LLM) ตระกูล Qwen (27B+) บนขุมพลัง AMD MI300X ผ่านระบบ vLLM Monster Engine… See the full description on the dataset page: https://huggingface.co/datasets/Phonsiri/legal-chat-sft-dataset.texttext-generation10K<n<100K1 likes45 downloads5mo agoHugging Face04SerFabio89 /italian-open-sft-chat-dataset Italian Open SFT Chat Dataset An Italian-first, model-neutral synthetic SFT and chat dataset for fine-tuning Italian-capable LLMs. It targets instruction tuning, Italian chat behavior, structured output generation, JSON/YAML/CSV format following, coding assistance, safety refusals, multi-turn dialogue and reasoning-style final answers. This v0.1.0 package does not include long-context QA records. This dataset is intended for users searching for an Italian instruction tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/SerFabio89/italian-open-sft-chat-dataset.texttext-generation10K<n<100K0 likes43 downloads5mo agoHugging Face05Daffaadityp /smkn26-chat-sft-id-en SMKN 26 Jakarta — PEFT Chat Dataset Dataset percakapan bilingual (Indonesia + English) untuk fine-tuning asisten resmi SMK Negeri 26 Jakarta memakai PEFT (LoRA / DoRA / QLoRA). Dirancang agar model: menjawab fakta sekolah dengan akurat menangani pertanyaan orang tua / calon siswa secara natural menolak mengarang saat data tidak tersedia (anti-hallucination) Source of truth: Dataset SMKN 26 Jakarta.pdf → curated KB kb/smkn26_facts.jsonRebuild: python scripts/build_dataset.py… See the full description on the dataset page: https://huggingface.co/datasets/Daffaadityp/smkn26-chat-sft-id-en.texttext-generation1K<n<10K0 likes29 downloads2mo agoHugging Face06PoSTMEDIA /rosetta-ko-chat-synth-sftgated rosetta-ko-chat-synth-sft Korean-native general-purpose conversational data — diverse everyday instructions with natural Korean answers. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Chat Suite Sibling datasets from the same pipeline (each a separate repo): repo format rosetta-ko-chat-synth-sft (this repo) supervised fine-tuning (user/assistant messages)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-chat-synth-sft.texttext-generation100K<n<1M0 likes24 downloads14d agoHugging Face07CausalLM /Retrieval-SFT-Chatgated Retrieval-Based Multi-Turn Chat SFT Synthetic Data A year ago, we released CausalLM/Refined-Anime-Text, a thematic subset of a dataset generated using the then state-of-the-art LLMs. This dataset comprises 1 million entries synthesized through long-context models that rewrote multi-document web text inputs, intended for continued pre-training. We are pleased to note that this data has been employed in various training scenarios and in studies concerning data and internet culture. In… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/Retrieval-SFT-Chat.textquestion-answering100K<n<1M61 likes20 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.