datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
github-codeThe GitHub Code dataest consists of 115M code files from GitHub in 32 programming languages with 60 extensions totalling in 1TB of text data. The dataset was created from the GitHub dataset on BiqQuery.qwen35-4b-drpo-vs0f49th-trainer-logprobs
Qwen3.5 4B DRPO trainer logprobs from W&B run vs0f49th
This dataset contains the raw trainer-logprob JSONL shards saved by W&B run ai2-llm/open_instruct_internal/vs0f49th (qwen35_4b_drpo__42__1782345587).
Contents
Source run: https://wandb.ai/ai2-llm/open_instruct_internal/runs/vs0f49th
Source path: /weka/oe-adapt-default/allennlp/deletable_rollouts/
Filename pattern: qwen35_4b_drpo__42__1782345587_trainer_logprobs_step*_rank*.jsonl
Files: 4320 JSONL shards… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/qwen35-4b-drpo-vs0f49th-trainer-logprobs.ROCStories
ROCStories (prompt / continuation)
A reformatted version of the ROCStories
corpus, suitable for open-ended story-generation exercises.
Each example is a 5-sentence ROCStory, split into:
field
description
prompt
the first sentence of the story
continuation
the remaining four sentences
text
the full (unmodified) 5-sentence story
Splits
split
rows
train
70,676
validation
7,852
test
19,633
The train / validation split is a 90/10… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/ROCStories.baligh-hamasa
Balīgh
Instruction-tuning data for classical Arabic, built from printed books of the
Arabic philological tradition. 40,145 records across two books.
taḥwīl (تحويل) — rewrite modern Arabic, MSA or dialect, into classical Arabic.
qa — one linguistic fact from the book, asked in a real voice (student,
reader, writer, teacher, editor, preacher, learner) and answered in the
author's words.
sharḥ_kāmil — a verse explained under fixed scholarly headings from
several of its claims at… See the full description on the dataset page: https://huggingface.co/datasets/omarabb315/baligh-hamasa.tunisian-derja-unified-raw-corpusTunisian Derja Unified Raw Corpus
Dataset Description
Repository: hamzabouajila/tunisian-derja-unified-raw-corpus
Paper: Not yet published; dataset card serves as primary documentation
Point of Contact: Hamza Bouajila
License: CC-BY-SA-4.0
Dataset Summary
The Tunisian Derja Unified Raw Corpus is a comprehensive collection of ~802,659 text examples in Tunisian Arabic (Derja), a low-resource dialect of Arabic widely spoken in Tunisia. This raw corpus aggregates data from multiple sources… See the full description on the dataset page: https://huggingface.co/datasets/hamzabouajila/tunisian-derja-unified-raw-corpus.Hala-4.6M-SFT
Hala: Arabic-Centric Instruction & Translation Dataset
Paper: Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale
Authors: Hasan Abed Al Kader Hammoud*, Mohammad Zbeeb*, Bernard Ghanem
Affiliation: King Abdullah University of Science and Technology (KAUST)
*Equal contribution
In Arabic, حلا (Hala) conveys sweetness and beauty—qualities long associated with the language itself. In this spirit, we extend Hala to datasets that aim to enrich… See the full description on the dataset page: https://huggingface.co/datasets/hammh0a/Hala-4.6M-SFT.Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1
Persian Civil Procedure QA Dataset
Dataset Description
این مجموعهداده شامل پرسشوپاسخهای حقوقی به زبان فارسی در حوزه آیین دادرسی مدنی است.
هر نمونه شامل سه فیلد اصلی است:
question: پرسش حقوقی
answer: پاسخ پرسش
evidence_quote: عبارت دقیق و مستند از دادهٔ منبع که پاسخ بر اساس آن استخراج شده است
هدف مجموعهداده، فراهمکردن دادهای ساختاریافته برای آموزش، ارزیابی و توسعه مدلهای زبانی فارسی در زمینه پرسشوپاسخ حقوقی است.
Dataset Structure
نمونهای… See the full description on the dataset page: https://huggingface.co/datasets/hamidsalimi/Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1.UltraChatTR_50k
💬 UltraChat 50K – Türkçe Diyalog Veri Seti
UltraChat 50K, orijinal UltraChat veri setinden türetilmiş,55.046 Türkçe diyalog örneği içeren açık kaynak bir veri setidir.Veri, büyük dil modellerinin Türkçe konuşma anlayışı ve cevap kalitesini geliştirmek içinfine-tuning (SFT) amacıyla düzenlenmiştir.
📘 Veri Künyesi
Özellik
Açıklama
Toplam Satır Sayısı
55.046
Veri Formatı
JSON Lines, Parquet
Alanlar
instruction, input, output
Dil
Türkçe 🇹🇷
Lisans
MIT… See the full description on the dataset page: https://huggingface.co/datasets/hamuz/UltraChatTR_50k.A11y-CUA
A11y-CUA Dataset
A11y-CUA is a multimodal desktop interaction dataset for accessibility-focused computer-use agent research. It contains real task trajectories recorded on Windows across two human user groups and two computer use agents (CUAs), each operating under standard and accessibility-specific conditions. Every session captures: timestamped keyboard and mouse events, browser interaction logs, accessibility trees, screen video, and system audio.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/hamid735/A11y-CUA.magpie-qwen-turbo-27k
Magpie-Qwen-Turbo-27k
Aratako/Magpie-Tanuki-8B-annotated-96k
のアノテーションを利用して件数を減らし、outputをqwen-2.5-turboで再生成したSFT用の26728件のサブセットです。
用途
コーディングを除く小規模な日本語チャット用LLMのためのファインチューニングを想定しています。
使用データ
以下の条件で抽出したinstructionデータを利用して生成しました。
input_quality(クエリの質):excellent のみ
difficulty(難易度):very easy/easy/medium/hard
primary_tag(カテゴリ):
"Information seeking", # ユーザーがさまざまなトピックに関する特定の情報や事実を求めるクエリ。
"Reasoning", # 論理的思考、問題解決、または複雑なアイデアの処理が必要なクエリ。
"Planning", #… See the full description on the dataset page: https://huggingface.co/datasets/hama-jp/magpie-qwen-turbo-27k.Teeet
Demet Turkish Chat Dataset
Bu dataset, Türkçe sohbet modeli fine-tuning için hazırlanmış konuşma örnekleri içerir.
Dataset Bilgileri
Karakter: Demet
Yaş: 17
Şehir: Ankara
Dil: Türkçe
Format: ChatML (messages formatı)
Kullanım
from datasets import load_dataset
dataset = load_dataset("hammur/Teeet")
Format
Her örnek şu formatta:
{
"messages": [
{"role": "system", "content": "..."},
{"role": "user", "content": "..."},
{"role":… See the full description on the dataset page: https://huggingface.co/datasets/hammur/Teeet.haman-fa-wikipedia-articles-186k
Haman Persian Wikipedia Articles 186K
A prepared Persian article dataset used to train
Haman Persian Article Graph-LLM 125M.
Dataset Summary
The dataset is derived from the Persian Wikipedia pages-articles dump and was
cleaned and prepared using the native article workflow in the
Rakhshai Graph-based NLP project.
Language: Persian
Accepted articles: 185,906
Training records: 176,611
Validation records: 9,295
Validation ratio: 5%
Split seed: 42
Training format:… See the full description on the dataset page: https://huggingface.co/datasets/aria-haman/haman-fa-wikipedia-articles-186k.reddit-dataseturdu-emergency-calls
Urdu Emergency Call Conversations Dataset (Pakistan)
Overview
This dataset contains 5,000 curated Urdu emergency call conversation samples from the Pakistan region, designed to support training and evaluation of Urdu Large Language Models (LLMs) for emergency response, command centers, and interpreter-style systems.
The conversations simulate real-world emergency scenarios such as:
Floods
Medical emergencies
Accidents
Crimes
Natural disasters
Public safety incidents
The… See the full description on the dataset page: https://huggingface.co/datasets/hamza-amin/urdu-emergency-calls.HBK08-subtitles
HBK08-subtitles
2024-05-26之前红警HBK08所有视频的数据集,数据来源于网络爬虫
metadata.tsv: 视频元数据: 包括url,bvid, UP主,标题,播放量,日期,时长;
raw_data.json: 原始视频字幕信息
text_cut.json: 文本分割标注,标记了充电投币感谢,以及视频中出现的广告
<begin>: 正文开始
<ad_begin>: 广告开始
<ad_end>: 广告结束
[discarded]: 标记这个文档被丢弃
ad_key_words.txt: 广告关键词
corrected_data.tsv: 粗略清洗的文本数据
主要采用文本替换+少量人工校对替换错误的字幕
少量以[verified]标签开头的经过了人工听写校对
3gpp_datasethamma-data
HAMMA DevOps Failure States Dataset ⬛⬜
This dataset was created to fine-tune the core intelligence engine for HAMMA — a local-first, zero-cloud SSH client built for the Gemma 4 Good Hackathon. It contains 3,701 curated problem-solution pairs covering the full surface area of Linux server administration: permissions, systemd, networking, Docker, Kubernetes, databases, storage, security, and CI/CD.
🎯 The "Zero-Fluff" Philosophy
Standard instruction-tuned LLMs respond… See the full description on the dataset page: https://huggingface.co/datasets/xayrullonematov/hamma-data.beavertails_with_refusals_trainThis dataset is associated with the research presented in the paper Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks.
The paper proposes Patcher, a method inspired by adversarial training and bi-level optimization, to combat full-parameter malicious finetuning attacks on large language models (LLMs).
Links
Paper: https://huggingface.co/papers/2606.07970
GitHub Repository: https://github.com/haomingwen/patcher
Data Format
According… See the full description on the dataset page: https://huggingface.co/datasets/Hammington/beavertails_with_refusals_train.pakistan-political-leaders-chatml-dataset
🇵🇰 Pakistan Political Leaders ChatML Dataset
🚀 A high-quality instruction-style dataset designed for fine-tuning Large Language Models (LLMs) on Pakistani political history and leadership.
This dataset contains approximately 2500 curated question-answer pairs in ChatML format, enabling models to understand and respond to queries about major political figures in Pakistan.
🎯 Objective
The goal of this dataset is to:
Train LLMs to act as a knowledgeable political… See the full description on the dataset page: https://huggingface.co/datasets/Hamzasajjad38/pakistan-political-leaders-chatml-dataset.hambobos-wikipedia-dataset_100K
Напоминую что оригинал есть на аккаунте hambobo14.
License
The dataset content is derived from Wikipediaand distributed under the CC BY-SA 3.0 license.This dataset release itself is shared under the CC-BY-NC-2.0 for metadata and structure.
MixInstruct-Llama-3This dataset contains responses to 5000 questions from the test split of MixInstruct dataset by Llama models, including:
Llama 3 Instruct
meta-llama-3-8b-instruct
meta-llama-3-70b-instruct
Llama 2 Chat
llama-2-7b-chat
llama-2-70b-chat
This dataset can be used as a starting point for evaluating the performance of Llama models.
Data Format
The dataset is stored as a .jsonl file with each object containting two fields: (1) model storing which LLM generated… See the full description on the dataset page: https://huggingface.co/datasets/hamnaanaa/MixInstruct-Llama-3.rus-Wikipedia-Dataset-1K
🇷🇺 rus-Wikipedia-Dataset-1K
Мини-датасет из 1000 статей русской Википедии для тестов NLP.
Файл датасета - Dataset.jsonl
📚 Формат
content — содержимое статьи
title — заголовок
⚙️ Использование
# -- Импорт...
from datasets import load_dataset
# -- Загружаем датасет на переменную data...
data = load_dataset("hambobo14/rus-Wikipedia-Dataset-1K")
# -- Проверяем...
print(data["content"][0])
# Готово!
📜 License
The dataset content is derived… See the full description on the dataset page: https://huggingface.co/datasets/hambobo14/rus-Wikipedia-Dataset-1K.beavertails_330k
Beavertails with Refusals Train
This dataset is used for alignment training to defend against malicious finetuning, as presented in the paper Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks.
Project Resources
Paper: https://huggingface.co/papers/2606.07970
Repository: https://github.com/haomingwen/patcher
Dataset Description
This dataset consists of prompts and safety-aligned responses (refusals) used to train… See the full description on the dataset page: https://huggingface.co/datasets/Hammington/beavertails_330k.hambobos-wikipedia-dataset_100K
This is finally released!
size 110K
ru and en
eazy for use
License
The dataset content is derived from Wikipediaand distributed under the CC BY-SA 3.0 license.This dataset release itself is shared under the MIT License for metadata and structure.
3gpp_fine-tunning-dataCoq-Hammer
Coq-Hammer
Structured dataset from CoqHammer — Automation for dependent type theory via ATPs.
Source
Repository: https://github.com/lukaszcz/coqhammer
Commit: 9e081180c6b00ca3925cf84b08d8621f562f8285
Files: 28
License: lgpl-2.1
Schema
Column
Type
Description
statement
string
Declaration signature/claim with the leading keyword removed (verbatim slice); the full declaration minus its proof
proof
string
Verbatim proof/body, empty if the… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/Coq-Hammer.HAMLET
HAMLET: A Hierarchical and Adaptive Multi-Agent Framework for Live Embodied Theatrics
Official dataset for HAMLET
We proposed a multi-agent framework HAMLET that decouples offline planning and online performance in AI theatrics scenarios. HAMLET excels in creating expressive, coherent, real-time and physically interactive drama experiences in a fully autonomous manner.
English | 简体中文
📖 Overview
🎉 News
[2026.03.19]… See the full description on the dataset page: https://huggingface.co/datasets/Tsumugii/HAMLET.data-science-chatbot
📊 Data Science Chatbot Dataset (2000 Samples)
🚀 A high-quality instruction-style dataset designed for fine-tuning Large Language Models (LLMs) on Data Science concepts.
This dataset contains ~2000 curated question-answer pairs in ChatML format, enabling models to learn how to explain, define, and discuss core data science topics in a clear and beginner-friendly way.
🎯 Objective
The goal of this dataset is to:
Train LLMs to act as a Data Science Tutor
Provide clear… See the full description on the dataset page: https://huggingface.co/datasets/Hamzasajjad38/data-science-chatbot.Israel-HAMAS_war_news
Dataset Card for Israel-HAMAS war news
Dataset Description
Point of Contact: Alexander Akhterov
Dataset Summary
The "Israel-HAMAS war news" dataset is an English-language dataset of news about Israel war against the terrorist organization -
HAMAS that happened after "black Saturday" - massive murders of civilian Israeli people on the 7th of October 2023.
We've accumulated news from the following sources:
BBC (live news) - from 2023-11-05 to 2023-11-18. Total:… See the full description on the dataset page: https://huggingface.co/datasets/aav-ds/Israel-HAMAS_war_news.Nexus-AI
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Hamiline/Nexus-AI.
