CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01oi-uae /cyber-securitygated Cybersecurity Instruction-Tuning Dataset A large, cleaned, multi-domain cybersecurity chat dataset for LLM finetuning, built from 198 distinct sources spanning offensive security, blue-team operations, vulnerability intelligence, cloud/AWS security, malware analysis, digital forensics, and more. Every record is normalized to the standard messages chat format and deduplicated at both file and record level. ⚠️ Research use only. This dataset is provided exclusively for… See the full description on the dataset page: https://huggingface.co/datasets/oi-uae/cyber-security.textquestion-answering1M<n<10M20 likes495 downloads14d agoHugging Face02makesi1 /service-uahj 传奇私服分布式路由与自动化接口索引库 - Batch 001 本项目托管了用于全球多元化检索与高可用分布式网络路由数据(核心挂载:传奇私服)。 📂 区域节点集群子目录 (Spider Pool Indexes) 👉 新开传奇私服 - 传奇私服网 - 传奇私服发布 —— 承载源站 🌐 api.agergames.com 👉 传奇私服推荐 - 新开传奇私服 - 传奇私服 —— 承载源站 🌐 api.amoygame.com 👉 今日传奇私服 - 传奇私服推荐 - 传奇私服网 —— 承载源站 🌐 api.gamecctv.com 👉 传奇私服999 - 传奇私服发布网 - 传奇私服 —— 承载源站 🌐 api.gamerling.com 👉 传奇私服发布 - 传奇私服999 - 新开传奇私服 —— 承载源站 🌐 api.games-vr.com 👉 传奇私服 - 今日传奇私服 - 传奇私服999 —— 承载源站 🌐 api.j5games.com 👉 今日传奇私服 - 传奇私服发布 -… See the full description on the dataset page: https://huggingface.co/datasets/makesi1/service-uahj.text-generation0 likes227 downloads2mo agoHugging Face03UAzimov /uzbek-instruct-llmuzbek-instruct-llm is a corpus of more than 15,000 records. It's made for instruct fine-tuning large language models for Uzbek language. It's mostly translated from other instruct datasets with some extra data added texttext-generation10K<n<100K9 likes158 downloads2y agoHugging Face04MariaOnyshchuk /ua-llm-router-eval UA Specialist Router — evaluation progress Score tables, routing stats, and run metadata from the diploma project MariaOnyshchuk/ua-llm-router: a rules-based router over open Ukrainian specialists (Mamay-4B, Lapa-12B, Aya Expanse 8B, Qwen-Coder). This dataset is the progress log of pinned JSON summaries, not a dump of every generation. What is included Path Contents progress_ledger.csv Flattened metric rows across weeks (best table for browsing)… See the full description on the dataset page: https://huggingface.co/datasets/MariaOnyshchuk/ua-llm-router-eval.text-generationn<1K0 likes135 downloads11d agoHugging Face05robinhad /tiny-ua-bench-responses Tiny-UA-Bench Responses This dataset contains the response matrix for Tiny-UA-Bench. The matrix contains 919,160 model and item records. The matrix covers 20 models and 45,958 items. The evaluation excludes FLORES and LongFLORES. Use Use this dataset to reproduce the benchmark compression analysis. Do not use a held-out model response to fit a selector or predictor. Use the reference and held-out split definitions from the code repository. Load the data with the… See the full description on the dataset page: https://huggingface.co/datasets/robinhad/tiny-ua-bench-responses.tabulartext-generation100K<n<1M0 likes100 downloads26d agoHugging Face06uaytug /UCDS uCoder Dataset A high-quality, deduplicated dataset for training coding and mathematics language models. Dataset Statistics Total Samples: 420,686 Format: ChatML (messages array) Languages: Python, JavaScript, C++, Java, and more Sources This dataset merges and cleans data from: Source Samples Description ByteDance-Seed/Code-Contests-Plus 10,293 Competitive programming open-r1/codeforces 47,558 Codeforces problems sahil2801/CodeAlpaca-20k 15… See the full description on the dataset page: https://huggingface.co/datasets/uaytug/UCDS.text-generation100K<n<1M0 likes98 downloads9mo agoHugging Face07uaytug /fumea-dataset FUMEA Dataset FUMEA-Dataset is a merged, curated, and deduplicated corpus designed for Supervised Fine-Tuning (SFT) of large language models. It unifies two specialized domains — tool-use / function-calling and financial analysis — into a single, training-ready resource. All samples are pre-formatted with the Qwen3 chat template (<|im_start|> / <|im_end|>) and require no additional preprocessing. This dataset is the primary training resource behind the FUMEA-F model family, which… See the full description on the dataset page: https://huggingface.co/datasets/uaytug/fumea-dataset.texttext-generation100K<n<1M0 likes53 downloads7mo agoHugging Face08hausmer /ua-tg-misc UA/TG misc — 15 small channel stubs A bundle of the 15 small Telegram channel exports (10 of them kept non-empty text rows; the rest exported only media-only posts, stripped here) — stubs and partially scraped channels that did not reach the full 10k-message scrape that the main corpora (hausmer/ukr-tg-media, hausmer/ukr-tg-satire) got. Most are 1–30 text posts (some channels were only reachable briefly, others export empty media rows). This corpus contains 67 text posts across… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/ua-tg-misc.texttext-generationn<1K0 likes45 downloads10d agoHugging Face09CJJones /Synthetic_UAV_Scenario_LLM_MultiTurn_GPS_Navigation Dataset Card for CJJones/Synthetic_UAV_Scenario_LLM_MultiTurn_GPS_Navigation The full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com Want more? 🚀 Get the AI Startup Bundle from Gumroad. Dataset Summary This dataset contains synthetic multi-turn UAV (Unmanned Aerial Vehicle) flight scenarios with realistic GPS navigation challenges, flight mode transitions, and system diagnostics. The scenarios simulate various flight conditions… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Synthetic_UAV_Scenario_LLM_MultiTurn_GPS_Navigation.texttext-generationn<1K0 likes41 downloads7mo agoHugging Face10cidtd-mod-ua /WizardLM-ukrainian WizardLM Translated to Ukrainian 🇺🇦 Dataset Description A Ukrainian language dataset comprising 140,000+ records translated from the WizardLM dataset. This dataset is suitable for various natural language processing tasks. This is not merged with original ShareGPT threads. Data translated via using Google Gemini Pro API. Слава Україні! Disclaimer Prepare data before your usage. There are some errors in texts, so be carefull. How to Use This… See the full description on the dataset page: https://huggingface.co/datasets/cidtd-mod-ua/WizardLM-ukrainian.texttext-generation100K<n<1M2 likes39 downloads3y agoHugging Face11oi-uae /CVEsgated CVEs — a full-coverage CVE chat dataset 1,625,017 chat conversations covering all 361,190 usable CVEs (1999–2026), built for fine-tuning cybersecurity assistants. Every known CVE in the official CVE List with severity enrichment from NVD (via the fkie-cad community feeds), rendered as English user/assistant conversations with varied phrasings, honest handling of missing data, and a per-CVE 99/1 train/validation split with zero leakage. The schema matches oi-uae/cyber-security… See the full description on the dataset page: https://huggingface.co/datasets/oi-uae/CVEs.texttext-generation1M<n<10M2 likes34 downloads11d agoHugging Face12anon-researcher-ua /ua-codeforces-cots-open-r1-for-training Version of anon-researcher-ua/ua-codeforces-cots-open-r1 prepared for model training texttext-generation1K<n<10K0 likes29 downloads1y agoHugging Face13lopozz /UA4RAG-it UA4RAG 📘 Dataset Summary UA4RAG (UnAnswerable for RAG) is a collection of datasets designed to train and evaluate language models on generating and recognizing unanswerable factual questions and appropriate non-answers given a reference text. In retrieval-augmented generation (RAG) systems, retrieved contexts are often tangential to user queries. This dataset addresses the critical challenge of training models to recognize when sufficient evidence is absent and to… See the full description on the dataset page: https://huggingface.co/datasets/lopozz/UA4RAG-it.texttext-generation1K<n<10K0 likes28 downloads11mo agoHugging Face14SalahALHaismawi /uae-laws-irac UAE Laws Q&A Dataset (IRAC Format) A high-quality dataset of 9,477 question-answer pairs about UAE laws, formatted in IRAC (Issue, Rule, Application, Conclusion) legal reasoning structure. Dataset Creation Source Documents The dataset was built from a comprehensive collection of UAE legal documents, including: Federal Decrees and Laws Cabinet Resolutions Ministerial Decisions Civil and Commercial Codes Labor Law Traffic Law And more Creation Process… See the full description on the dataset page: https://huggingface.co/datasets/SalahALHaismawi/uae-laws-irac.textquestion-answering1K<n<10K1 likes24 downloads8mo agoHugging Face15overthelex /ua-legal-citation-grounded-sft UA Legal Citation-Grounded SFT A supervised fine-tuning set of citation-grounded legal question-answering examples in Ukrainian. Every assistant answer attributes each factual claim to a specific source with a [doc:ID] marker that refers to a real court decision passage placed in the prompt. The set is built to train and study retrieval-grounded generation where faithfulness of citations, not just answer quality, is the target. How it was built Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/ua-legal-citation-grounded-sft.texttext-generation10K<n<100K0 likes23 downloads2mo agoHugging Face16obadabaq /structured-uae-laws Dataset Card for structured-uae-laws This dataset is a collection of question & answers about the laws and regulations in the United Arab Emirates. It covers different areas of law like: economy and business family and community finance and banking industry and technical standardisation justice and juiciary, labour residency and leberal professions security and safety tax Dataset Sources Repository Base Dataset United Arab Emirates Legislations… See the full description on the dataset page: https://huggingface.co/datasets/obadabaq/structured-uae-laws.textquestion-answering1K<n<10K1 likes22 downloads2y agoHugging Face17anon-researcher-ua /ua-codeforces-cots-open-r1 Dataset Summary ua-codeforces-cots-open-r1 is a Ukrainian-focused derivative of open-r1/codeforces-cots that: includes 1550 Python solutions from original dataset generated by DeepSeek-R1; adds Ukrainian translations of Codeforces task statements, I/O formats, notes, and editorials; provides Ukrainian translation of original ("high") reasoning obtained with DeepSeek-V3; adds “low” reasoning in Ukrainian by DeepSeek-R1 based on original reasoning and task statements; ships… See the full description on the dataset page: https://huggingface.co/datasets/anon-researcher-ua/ua-codeforces-cots-open-r1.tabulartext-generation1K<n<10K0 likes21 downloads1y agoHugging Face18nikes64 /ualpaca-gpt4 Dataset Card for "alpaca-gpt4-cleaned" This dataset contains Ukrainian Instruction-Following translated by facebook/nllb-200-3.3B The dataset was originaly shared in this repository: https://github.com/tloen/alpaca-lora Licensing Information The dataset is available under the Creative Commons NonCommercial (CC BY-NC 4.0). text-generation10K<n<100K2 likes20 downloads3y agoHugging Face19podarok /kobza-cleaned-ua kobza-cleaned-ua Cleaned Ukrainian-language subset of Goader/kobza dataset with Russian content filtered out. Dataset Details Dataset Description This dataset is a cleaned and filtered version of the Goader/kobza corpus, removing Russian language content to create a pure Ukrainian language dataset suitable for training language models. The original kobza dataset contains ~60B tokens across 97 million documents. This cleaned version maintains ~59B tokens (98.5%… See the full description on the dataset page: https://huggingface.co/datasets/podarok/kobza-cleaned-ua.texttext-generation1M<n<10M0 likes19 downloads9mo agoHugging Face20vikramlingam /UAE-corptax-training Dataset Card for UAE Corporate Tax Q&A Dataset Dataset Summary This dataset contains 1,283 instruction-response pairs covering UAE Corporate Tax regulations from 2022-2025. Built from several official sources. Each response includes proper legal citations. Perfect for fine-tuning LLMs for UAE tax advisory, building RAG systems, or training tax compliance tools. This is for educational purpose only. Supported Tasks Instruction Following: Train models to answer… See the full description on the dataset page: https://huggingface.co/datasets/vikramlingam/UAE-corptax-training.textquestion-answering1K<n<10K0 likes18 downloads9mo agoHugging Face21shamotskyi /ua_cbt_stories Dataset Card for UA-CBT Stories This dataset was generated in the context of Eval-UA-tion 1.0 benchmark for evaluating Ukrainian language models (paper, thesis). It contains the generated and manually corrected stories used for the UA-CBT (Ukrainian Children's Book Test) task. The dataset contains Ukrainian-language stories, LLM-generated in multiple steps and then manually corrected (or marked as unusable if fixing them was too hard). For each story, the original LLM prompt, all… See the full description on the dataset page: https://huggingface.co/datasets/shamotskyi/ua_cbt_stories.tabulartext-generationn<1K0 likes17 downloads2y agoHugging Face22alexshynkarenk0 /UA-Safety-Align-Sample UA-Safety-Align: Ukrainian Red Teaming & Safety Dataset (Sample) Developed by DavidLab Contact: founder@davidlab.techStatus: Sample (50 examples). Full dataset (4,800+ examples) available for commercial licensing. 🚀 Dataset Summary UA-Safety-Align is a specialized dataset designed for Red Teaming, Safety Alignment, and Robustness Testing of Large Language Models (LLMs) specifically in the Ukrainian language context. While most safety datasets focus on English… See the full description on the dataset page: https://huggingface.co/datasets/alexshynkarenk0/UA-Safety-Align-Sample.texttext-generationn<1K1 likes16 downloads8mo agoHugging Face23obadabaq /uae-laws Dataset Card for UAE-Laws This dataset is a collection of information about the laws and regulations in the United Arab Emirates. It covers different areas of law like: economy and business family and community finance and banking industry and technical standardisation justice and juiciary, labour residency and leberal professions security and safety tax Dataset Sources United Arab Emirates Legislations Dataset Structure The ./uae-laws.csv… See the full description on the dataset page: https://huggingface.co/datasets/obadabaq/uae-laws.textquestion-answering1K<n<10K4 likes15 downloads2y agoHugging Face24robinhad /UAlpaca2.0 UAlpaca 2.0 textsummarization10K<n<100K0 likes13 downloads2y agoHugging Face25adarshrajesh /uae-adab-tutor-600 UAE Adab Tutor 600 This is the 600-conversation supervised fine-tuning dataset used for adarshrajesh/uae-adab-tutor-qwen3-4b. Release version: exact-silver v1 Complete-600. Behavior spec Across a pressured multi-turn lesson, the tutor should teach the academic content accurately, correct the specific work without humiliating the learner, protect learner authorship and assessment integrity, allow respectful evidence-based disagreement with adults, avoid religious… See the full description on the dataset page: https://huggingface.co/datasets/adarshrajesh/uae-adab-tutor-600.texttext-generationn<1K0 likes13 downloads3mo agoHugging Face26oshyshatskyi /ua-council-decisions Ukrainian Municipal Council Decisions — Masthead Identity Extraction Structured-extraction dataset of 1,075 Ukrainian municipal council decisions (рішення) from 43 local councils (громади / ради) — balanced to exactly 25 decisions per council, each paired with the five identity fields that appear in the document masthead. The task: given the full text of a single decision, extract its masthead identity. These are public government records. All personal names in the data are… See the full description on the dataset page: https://huggingface.co/datasets/oshyshatskyi/ua-council-decisions.tabulartext-generation1K<n<10K0 likes11 downloads2mo agoHugging Face27NLPForUA /ua-code-benchgated LLM Code Generation Benchmark for Ukrainian language Preprint: https://arxiv.org/pdf/2511.05040 Updates 17/10/2025: paper presented at "Informatics. Culture. Technology" conference; 18/09/2025: added data preparation and evaluation notebooks (check notebooks readme first); 17/09/2025: updated result chart; added gpt-5, gpt-oss, and grok-4 evaluations. Thousands of programming tasks in Ukrainian language combined with graded Python solutions (code… See the full description on the dataset page: https://huggingface.co/datasets/NLPForUA/ua-code-bench.tabulartext-generation1K<n<10K0 likes10 downloads2mo agoHugging Face28andrewwe /ua-toxic-light Ukrainian Style Chat Mix Chat-format dataset for Ukrainian style adaptation. Splits train: 5820 validation: 90 test: 90 Schema Each row has: messages: list of chat turns (role, content) source: source dataset id optional style_toxic: 0/1 style marker Notes Intended for controlled style tuning. Keep style data as a minority share during model training. texttext-generation1K<n<10K0 likes10 downloads6mo agoHugging Face29anon-researcher-ua /ua-code-benchgated LLM Code Generation Benchmark for Ukrainian language Preprint: https://arxiv.org/pdf/2511.05040 Updates 18/09/2025: added data preparation and evaluation notebooks (check notebooks readme first); 17/09/2025: updated result chart; added gpt-5, gpt-oss, and grok-4 evaluations; 17/09/2025: paper - end of September, stay tuned; Thousands of programming tasks in Ukrainian language combined with graded Python solutions (code + reasoning) by leading LLMs (DeepSeek… See the full description on the dataset page: https://huggingface.co/datasets/anon-researcher-ua/ua-code-bench.tabulartext-generation1K<n<10K0 likes7 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.