CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Hammad47 /github-codeThe GitHub Code dataest consists of 115M code files from GitHub in 32 programming languages with 60 extensions totalling in 1TB of text data. The dataset was created from the GitHub dataset on BiqQuery.text-generation0 likes425 downloads10mo agoHugging Face02hamishivi /qwen35-4b-drpo-vs0f49th-trainer-logprobs Qwen3.5 4B DRPO trainer logprobs from W&B run vs0f49th This dataset contains the raw trainer-logprob JSONL shards saved by W&B run ai2-llm/open_instruct_internal/vs0f49th (qwen35_4b_drpo__42__1782345587). Contents Source run: https://wandb.ai/ai2-llm/open_instruct_internal/runs/vs0f49th Source path: /weka/oe-adapt-default/allennlp/deletable_rollouts/ Filename pattern: qwen35_4b_drpo__42__1782345587_trainer_logprobs_step*_rank*.jsonl Files: 4320 JSONL shards… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/qwen35-4b-drpo-vs0f49th-trainer-logprobs.tabulartext-generation10K<n<100K0 likes299 downloads3mo agoHugging Face03hamishivi /ROCStories ROCStories (prompt / continuation) A reformatted version of the ROCStories corpus, suitable for open-ended story-generation exercises. Each example is a 5-sentence ROCStory, split into: field description prompt the first sentence of the story continuation the remaining four sentences text the full (unmodified) 5-sentence story Splits split rows train 70,676 validation 7,852 test 19,633 The train / validation split is a 90/10… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/ROCStories.texttext-generation10K<n<100K1 likes224 downloads5mo agoHugging Face04omarabb315 /baligh-hamasa Balīgh Instruction-tuning data for classical Arabic, built from printed books of the Arabic philological tradition. 40,145 records across two books. taḥwīl (تحويل) — rewrite modern Arabic, MSA or dialect, into classical Arabic. qa — one linguistic fact from the book, asked in a real voice (student, reader, writer, teacher, editor, preacher, learner) and answered in the author's words. sharḥ_kāmil — a verse explained under fixed scholarly headings from several of its claims at… See the full description on the dataset page: https://huggingface.co/datasets/omarabb315/baligh-hamasa.texttext-generation10K<n<100K0 likes133 downloads22d agoHugging Face05hamzabouajila /tunisian-derja-unified-raw-corpusTunisian Derja Unified Raw Corpus Dataset Description Repository: hamzabouajila/tunisian-derja-unified-raw-corpus Paper: Not yet published; dataset card serves as primary documentation Point of Contact: Hamza Bouajila License: CC-BY-SA-4.0 Dataset Summary The Tunisian Derja Unified Raw Corpus is a comprehensive collection of ~802,659 text examples in Tunisian Arabic (Derja), a low-resource dialect of Arabic widely spoken in Tunisia. This raw corpus aggregates data from multiple sources… See the full description on the dataset page: https://huggingface.co/datasets/hamzabouajila/tunisian-derja-unified-raw-corpus.texttext-generation100K<n<1M0 likes114 downloads1y agoHugging Face06hammh0a /Hala-4.6M-SFT Hala: Arabic-Centric Instruction & Translation Dataset Paper: Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale Authors: Hasan Abed Al Kader Hammoud*, Mohammad Zbeeb*, Bernard Ghanem Affiliation: King Abdullah University of Science and Technology (KAUST) *Equal contribution In Arabic, حلا (Hala) conveys sweetness and beauty—qualities long associated with the language itself. In this spirit, we extend Hala to datasets that aim to enrich… See the full description on the dataset page: https://huggingface.co/datasets/hammh0a/Hala-4.6M-SFT.texttext-generation1M<n<10M5 likes102 downloads1y agoHugging Face07hamidsalimi /Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1 Persian Civil Procedure QA Dataset Dataset Description این مجموعه‌داده شامل پرسش‌وپاسخ‌های حقوقی به زبان فارسی در حوزه آیین دادرسی مدنی است. هر نمونه شامل سه فیلد اصلی است: question: پرسش حقوقی answer: پاسخ پرسش evidence_quote: عبارت دقیق و مستند از دادهٔ منبع که پاسخ بر اساس آن استخراج شده است هدف مجموعه‌داده، فراهم‌کردن داده‌ای ساختاریافته برای آموزش، ارزیابی و توسعه مدل‌های زبانی فارسی در زمینه پرسش‌وپاسخ حقوقی است. Dataset Structure نمونه‌ای… See the full description on the dataset page: https://huggingface.co/datasets/hamidsalimi/Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1.textquestion-answeringn<1K1 likes73 downloads2mo agoHugging Face08hamuz /UltraChatTR_50k 💬 UltraChat 50K – Türkçe Diyalog Veri Seti UltraChat 50K, orijinal UltraChat veri setinden türetilmiş,55.046 Türkçe diyalog örneği içeren açık kaynak bir veri setidir.Veri, büyük dil modellerinin Türkçe konuşma anlayışı ve cevap kalitesini geliştirmek içinfine-tuning (SFT) amacıyla düzenlenmiştir. 📘 Veri Künyesi Özellik Açıklama Toplam Satır Sayısı 55.046 Veri Formatı JSON Lines, Parquet Alanlar instruction, input, output Dil Türkçe 🇹🇷 Lisans MIT… See the full description on the dataset page: https://huggingface.co/datasets/hamuz/UltraChatTR_50k.texttext-generation10K<n<100K2 likes52 downloads11mo agoHugging Face09hamid735 /A11y-CUA A11y-CUA Dataset A11y-CUA is a multimodal desktop interaction dataset for accessibility-focused computer-use agent research. It contains real task trajectories recorded on Windows across two human user groups and two computer use agents (CUAs), each operating under standard and accessibility-specific conditions. Every session captures: timestamped keyboard and mouse events, browser interaction logs, accessibility trees, screen video, and system audio. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/hamid735/A11y-CUA.text-generation100K<n<1M0 likes47 downloads7d agoHugging Face10hama-jp /magpie-qwen-turbo-27k Magpie-Qwen-Turbo-27k Aratako/Magpie-Tanuki-8B-annotated-96k のアノテーションを利用して件数を減らし、outputをqwen-2.5-turboで再生成したSFT用の26728件のサブセットです。 用途 コーディングを除く小規模な日本語チャット用LLMのためのファインチューニングを想定しています。 使用データ 以下の条件で抽出したinstructionデータを利用して生成しました。 input_quality(クエリの質):excellent のみ difficulty(難易度):very easy/easy/medium/hard primary_tag(カテゴリ): "Information seeking", # ユーザーがさまざまなトピックに関する特定の情報や事実を求めるクエリ。 "Reasoning", # 論理的思考、問題解決、または複雑なアイデアの処理が必要なクエリ。 "Planning", #… See the full description on the dataset page: https://huggingface.co/datasets/hama-jp/magpie-qwen-turbo-27k.texttext-generation10K<n<100K0 likes33 downloads2y agoHugging Face11hammur /Teeet Demet Turkish Chat Dataset Bu dataset, Türkçe sohbet modeli fine-tuning için hazırlanmış konuşma örnekleri içerir. Dataset Bilgileri Karakter: Demet Yaş: 17 Şehir: Ankara Dil: Türkçe Format: ChatML (messages formatı) Kullanım from datasets import load_dataset dataset = load_dataset("hammur/Teeet") Format Her örnek şu formatta: { "messages": [ {"role": "system", "content": "..."}, {"role": "user", "content": "..."}, {"role":… See the full description on the dataset page: https://huggingface.co/datasets/hammur/Teeet.texttext-generationn<1K0 likes30 downloads8mo agoHugging Face12aria-haman /haman-fa-wikipedia-articles-186k Haman Persian Wikipedia Articles 186K A prepared Persian article dataset used to train Haman Persian Article Graph-LLM 125M. Dataset Summary The dataset is derived from the Persian Wikipedia pages-articles dump and was cleaned and prepared using the native article workflow in the Rakhshai Graph-based NLP project. Language: Persian Accepted articles: 185,906 Training records: 176,611 Validation records: 9,295 Validation ratio: 5% Split seed: 42 Training format:… See the full description on the dataset page: https://huggingface.co/datasets/aria-haman/haman-fa-wikipedia-articles-186k.texttext-generation100K<n<1M1 likes29 downloads2mo agoHugging Face13hamzasibous /reddit-datasetimagetext-generation10K<n<100K0 likes28 downloads7mo agoHugging Face14hamza-amin /urdu-emergency-calls Urdu Emergency Call Conversations Dataset (Pakistan) Overview This dataset contains 5,000 curated Urdu emergency call conversation samples from the Pakistan region, designed to support training and evaluation of Urdu Large Language Models (LLMs) for emergency response, command centers, and interpreter-style systems. The conversations simulate real-world emergency scenarios such as: Floods Medical emergencies Accidents Crimes Natural disasters Public safety incidents The… See the full description on the dataset page: https://huggingface.co/datasets/hamza-amin/urdu-emergency-calls.texttext-generation1K<n<10K1 likes24 downloads9mo agoHugging Face15hammershock /HBK08-subtitles HBK08-subtitles 2024-05-26之前红警HBK08所有视频的数据集,数据来源于网络爬虫 metadata.tsv: 视频元数据: 包括url,bvid, UP主,标题,播放量,日期,时长; raw_data.json: 原始视频字幕信息 text_cut.json: 文本分割标注,标记了充电投币感谢,以及视频中出现的广告 <begin>: 正文开始 <ad_begin>: 广告开始 <ad_end>: 广告结束 [discarded]: 标记这个文档被丢弃 ad_key_words.txt: 广告关键词 corrected_data.tsv: 粗略清洗的文本数据 主要采用文本替换+少量人工校对替换错误的字幕 少量以[verified]标签开头的经过了人工听写校对 text-generation1K<n<10K1 likes23 downloads2y agoHugging Face16Hamza011 /3gpp_datasettext-generationn<1K0 likes21 downloads2y agoHugging Face17xayrullonematov /hamma-data HAMMA DevOps Failure States Dataset ⬛⬜ This dataset was created to fine-tune the core intelligence engine for HAMMA — a local-first, zero-cloud SSH client built for the Gemma 4 Good Hackathon. It contains 3,701 curated problem-solution pairs covering the full surface area of Linux server administration: permissions, systemd, networking, Docker, Kubernetes, databases, storage, security, and CI/CD. 🎯 The "Zero-Fluff" Philosophy Standard instruction-tuned LLMs respond… See the full description on the dataset page: https://huggingface.co/datasets/xayrullonematov/hamma-data.texttext-generation1K<n<10K1 likes21 downloads4mo agoHugging Face18Hammington /beavertails_with_refusals_trainThis dataset is associated with the research presented in the paper Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks. The paper proposes Patcher, a method inspired by adversarial training and bi-level optimization, to combat full-parameter malicious finetuning attacks on large language models (LLMs). Links Paper: https://huggingface.co/papers/2606.07970 GitHub Repository: https://github.com/haomingwen/patcher Data Format According… See the full description on the dataset page: https://huggingface.co/datasets/Hammington/beavertails_with_refusals_train.texttext-generation10K<n<100K0 likes20 downloads4mo agoHugging Face19Hamzasajjad38 /pakistan-political-leaders-chatml-dataset 🇵🇰 Pakistan Political Leaders ChatML Dataset 🚀 A high-quality instruction-style dataset designed for fine-tuning Large Language Models (LLMs) on Pakistani political history and leadership. This dataset contains approximately 2500 curated question-answer pairs in ChatML format, enabling models to understand and respond to queries about major political figures in Pakistan. 🎯 Objective The goal of this dataset is to: Train LLMs to act as a knowledgeable political… See the full description on the dataset page: https://huggingface.co/datasets/Hamzasajjad38/pakistan-political-leaders-chatml-dataset.texttext-generation1K<n<10K0 likes19 downloads5mo agoHugging Face20HBB-Community /hambobos-wikipedia-dataset_100K Напоминую что оригинал есть на аккаунте hambobo14. License The dataset content is derived from Wikipediaand distributed under the CC BY-SA 3.0 license.This dataset release itself is shared under the CC-BY-NC-2.0 for metadata and structure. text-generation0 likes18 downloads3mo agoHugging Face21hamnaanaa /MixInstruct-Llama-3This dataset contains responses to 5000 questions from the test split of MixInstruct dataset by Llama models, including: Llama 3 Instruct meta-llama-3-8b-instruct meta-llama-3-70b-instruct Llama 2 Chat llama-2-7b-chat llama-2-70b-chat This dataset can be used as a starting point for evaluating the performance of Llama models. Data Format The dataset is stored as a .jsonl file with each object containting two fields: (1) model storing which LLM generated… See the full description on the dataset page: https://huggingface.co/datasets/hamnaanaa/MixInstruct-Llama-3.texttext-generation10K<n<100K2 likes16 downloads2y agoHugging Face22hambobo14 /rus-Wikipedia-Dataset-1K 🇷🇺 rus-Wikipedia-Dataset-1K Мини-датасет из 1000 статей русской Википедии для тестов NLP. Файл датасета - Dataset.jsonl 📚 Формат content — содержимое статьи title — заголовок ⚙️ Использование # -- Импорт... from datasets import load_dataset # -- Загружаем датасет на переменную data... data = load_dataset("hambobo14/rus-Wikipedia-Dataset-1K") # -- Проверяем... print(data["content"][0]) # Готово! 📜 License The dataset content is derived… See the full description on the dataset page: https://huggingface.co/datasets/hambobo14/rus-Wikipedia-Dataset-1K.texttext-generation1K<n<10K1 likes16 downloads1y agoHugging Face23Hammington /beavertails_330k Beavertails with Refusals Train This dataset is used for alignment training to defend against malicious finetuning, as presented in the paper Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks. Project Resources Paper: https://huggingface.co/papers/2606.07970 Repository: https://github.com/haomingwen/patcher Dataset Description This dataset consists of prompts and safety-aligned responses (refusals) used to train… See the full description on the dataset page: https://huggingface.co/datasets/Hammington/beavertails_330k.texttext-generation100K<n<1M1 likes16 downloads4mo agoHugging Face24hambobo14 /hambobos-wikipedia-dataset_100K This is finally released! size 110K ru and en eazy for use License The dataset content is derived from Wikipediaand distributed under the CC BY-SA 3.0 license.This dataset release itself is shared under the MIT License for metadata and structure. text-generation100K<n<1M0 likes13 downloads6mo agoHugging Face25Hamza011 /3gpp_fine-tunning-datatexttext-generationn<1K5 likes12 downloads2y agoHugging Face26phanerozoic /Coq-Hammer Coq-Hammer Structured dataset from CoqHammer — Automation for dependent type theory via ATPs. Source Repository: https://github.com/lukaszcz/coqhammer Commit: 9e081180c6b00ca3925cf84b08d8621f562f8285 Files: 28 License: lgpl-2.1 Schema Column Type Description statement string Declaration signature/claim with the leading keyword removed (verbatim slice); the full declaration minus its proof proof string Verbatim proof/body, empty if the… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/Coq-Hammer.texttext-generationn<1K0 likes11 downloads4mo agoHugging Face27Tsumugii /HAMLET HAMLET: A Hierarchical and Adaptive Multi-Agent Framework for Live Embodied Theatrics Official dataset for HAMLET We proposed a multi-agent framework HAMLET that decouples offline planning and online performance in AI theatrics scenarios. HAMLET excels in creating expressive, coherent, real-time and physically interactive drama experiences in a fully autonomous manner. English | 简体中文 📖 Overview 🎉 News [2026.03.19]… See the full description on the dataset page: https://huggingface.co/datasets/Tsumugii/HAMLET.text-generation0 likes11 downloads6mo agoHugging Face28Hamzasajjad38 /data-science-chatbot 📊 Data Science Chatbot Dataset (2000 Samples) 🚀 A high-quality instruction-style dataset designed for fine-tuning Large Language Models (LLMs) on Data Science concepts. This dataset contains ~2000 curated question-answer pairs in ChatML format, enabling models to learn how to explain, define, and discuss core data science topics in a clear and beginner-friendly way. 🎯 Objective The goal of this dataset is to: Train LLMs to act as a Data Science Tutor Provide clear… See the full description on the dataset page: https://huggingface.co/datasets/Hamzasajjad38/data-science-chatbot.texttext-generation1K<n<10K0 likes11 downloads5mo agoHugging Face29aav-ds /Israel-HAMAS_war_news Dataset Card for Israel-HAMAS war news Dataset Description Point of Contact: Alexander Akhterov Dataset Summary The "Israel-HAMAS war news" dataset is an English-language dataset of news about Israel war against the terrorist organization - HAMAS that happened after "black Saturday" - massive murders of civilian Israeli people on the 7th of October 2023. We've accumulated news from the following sources: BBC (live news) - from 2023-11-05 to 2023-11-18. Total:… See the full description on the dataset page: https://huggingface.co/datasets/aav-ds/Israel-HAMAS_war_news.texttext-classification10K<n<100K2 likes10 downloads3y agoHugging Face30Hamiline /Nexus-AI Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Hamiline/Nexus-AI.texttext-generationn<1K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.