CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01its5Q /wikireading Dataset Card for Wikireading This is a dataset of book chapters scraped from a Russian website called Wikireading. Dataset Details Dataset Description Wikireading is a collection of non-fiction educational books in various domains: Biology, Art, History, Religion and much more. The books are highly educational and provide vast knowledge in different domains, making this dataset a good choice for pretraining. The resulting dataset contains ~26M rows, which in… See the full description on the dataset page: https://huggingface.co/datasets/its5Q/wikireading.texttext-generation1M<n<10M9 likes197 downloads2y agoHugging Face02benjaminmacklin /IT_Support_V2 Mack: IT Support & Admin Dataset 📋 Dataset Description This dataset consists of 100,000+ conversation logs focused on IT Support and IT Administration tasks. It was generated to fine-tune the "Mack" model—an AI persona designed to act as an expert Tier 1 & Tier 2 IT Helpdesk agent. The data covers a wide range of technical domains, including Windows troubleshooting, SQL Server administration, driver issues, network diagnostics, and hardware debugging. Curated by: [Dev… See the full description on the dataset page: https://huggingface.co/datasets/benjaminmacklin/IT_Support_V2.texttext-generation100K<n<1M2 likes121 downloads10mo agoHugging Face03itsgupta /proper-agents-data ProPer Agents — data Data for ProPer Agents: Proactivity Driven Personalized Agents for Advancing Knowledge Gap Navigation (ACL 2026). Paper · Adapters Three domains: code, medical, pwab (product recommendation). Layout {domain}/ raw/train.jsonl source examples raw/test.jsonl raw/{domain}_rga_{train,test}.jsonl RGA SFT data (Alpaca format) raw/{domain}_dga_{train,test}.jsonl DGA SFT data (Alpaca format)… See the full description on the dataset page: https://huggingface.co/datasets/itsgupta/proper-agents-data.texttext-generation1K<n<10K0 likes82 downloads2mo agoHugging Face04itsmebatuhan /bluesky-10m-posts-15-languages Dataset Card: Bluesky 10M Multilingual 📊 Overview Total Posts: 10,099,990 Languages: 15 (en, tr, es, pt, de, fr, ja, it, nl, pl, ru, ko, zh, ar, hi) Collection Period: August 9-12, 2026 Source: Bluesky Jetstream API (public firehose) Format: JSONL Size: ~3 GB 🌍 Language Distribution Language Code Posts % English en 6,843,995 67.8% Japanese ja 1,547,179 15.3% German de 373,626 3.7% Portuguese pt 331,093 3.3% Spanish es 325,865… See the full description on the dataset page: https://huggingface.co/datasets/itsmebatuhan/bluesky-10m-posts-15-languages.texttext-classification10M<n<100M0 likes64 downloads1mo agoHugging Face05its5Q /panorama Dataset Summary Dataset of satirical news from "Panorama", Russian "The Onion". Dataset Format Dataset is in JSONLines format, where "title" is the article title, and "body" are contents of the article. texttext-generation10K<n<100K6 likes59 downloads4y agoHugging Face06w1z4rd3k /it-support-l1-ticket-classification IT Support L1 Multilingual Dataset Dataset Summary IT Support L1 Multilingual Dataset is a synthetic enterprise help desk dataset for ticket classification and troubleshooting response generation. It contains realistic Level 1 IT support scenarios in English and Czech, designed for experiments in structured classification, response generation, and multilingual support workflow prototyping. This dataset contains synthetic IT Support L1 scenarios. The records were generated… See the full description on the dataset page: https://huggingface.co/datasets/w1z4rd3k/it-support-l1-ticket-classification.texttext-classificationn<1K0 likes36 downloads5mo agoHugging Face07itsluketwist /LangChoiceBenchgated LangChoiceBench Welcome to LangChoiceBench - the benchmark dataset for studying programming-language choice in reasoning LLMs. Introduced in the paper LangChoiceBench: Measuring and Explaining Programming-Language Choice in LLMs, the dataset covers 28 real-world software projects across 7 domains (mobile, frontend, low-latency, systems, embedded, games, enterprise) where Python is a known poor default, plus a 4-project python_control area where Python is the right choice… See the full description on the dataset page: https://huggingface.co/datasets/itsluketwist/LangChoiceBench.texttext-generationn<1K0 likes36 downloads7d agoHugging Face08itsalloverig /MIKE-dataset MIKE High-Signal Indian Legal Triage This is the curated instruction-tuning corpus for MIKE, an India-focused legal research and triage adapter. It contains 10,913 English examples designed for source-bounded reasoning, issue triage, document and evidence planning, structured output, and legacy/current criminal-law transition screening. Dataset composition 10,049 balanced, completion-deduplicated base examples; 244 explicit JSON-schema instruction variants; 500… See the full description on the dataset page: https://huggingface.co/datasets/itsalloverig/MIKE-dataset.texttext-generation10K<n<100K0 likes34 downloads3mo agoHugging Face09itsSHAS /clean_ukrainian-news Ukrainian News Dataset This is a dataset of news articles downloaded from various Ukrainian websites and Telegram channels. The dataset contains 22 567 099 JSON objects (news), total size ~67GB each with the following fields: title: The title of the news article text: The text of the news article, which may contain HTML tags(e.g., paragraphs, links, images, etc.) url: The URL of the news article datetime: The time of publication or when the article was parsed and added to… See the full description on the dataset page: https://huggingface.co/datasets/itsSHAS/clean_ukrainian-news.texttext-generation10M<n<100M0 likes30 downloads8mo agoHugging Face10ItsHotdogFred /kevin-v1-dataset Kevin V1 — NPC Conversation Dataset Synthetic player↔NPC conversations for training game NPC dialogue models. Generated with a 3-role pipeline (context / player / NPC) plus a judge that verifies every NPC reply is grounded (no hallucinated facts) and in-character. Format One conversation per row (JSON Lines). Each row: { "id": "conv_00042", "area_id": "01_emberpeak_forge", "npc": {"role": "blacksmith", "name": "...", "offers": [...], "knows_about": [...]}… See the full description on the dataset page: https://huggingface.co/datasets/ItsHotdogFred/kevin-v1-dataset.texttext-generationn<1K0 likes26 downloads4mo agoHugging Face11itsankitkp /swe-clarify SWE-Clarify-CFR: Game-Theoretic Clarification Dataset SWE-Clarify-CFR is a high-fidelity synthetic dataset designed to train Large Language Models (LLMs) to detect dangerous ambiguity in software engineering tasks. Standard LLMs suffer from "Helpfulness Bias", when presented with a vague request (e.g., "Flush the database"), they often guess the user's intent to be helpful. In high-stakes engineering, this can lead to catastrophic data loss or security breaches. This dataset solves… See the full description on the dataset page: https://huggingface.co/datasets/itsankitkp/swe-clarify.texttext-generation1K<n<10K0 likes24 downloads10mo agoHugging Face12its5Q /teletype Dataset Card for Teletype This dataset is a scrape of all articles published on teletype, a popular platform for publishing articles, especially in Telegram. The dataset includes the original article HTML, as well as text extracted using the trafilatura library with favor_recall=True and other metadata provided by teletype. Additionally, language identification was applied using the lingua-py library and the identification results are available in the lang column. Curated by: its5Q imagetext-generation1M<n<10M3 likes21 downloads2y agoHugging Face13itsZyn /ConvES30K ConvES10K-HQ-LLM — High-Quality Spanish Conversations (LLM-Generated) Recommended for AI training. This version replaces the template-based builds and is fully LLM-generated for true independence. Why this version Previous builds (30K and 10K-template) were template-based (gen.py + 178 situations): 13,156 distinct messages from 74,526 total → 82.3% duplicate messages, max x109 on narrative blocks (Mesa seis pegada a la mesa siete...) Coherence failures from… See the full description on the dataset page: https://huggingface.co/datasets/itsZyn/ConvES30K.texttext-generation10K<n<100K0 likes11 downloads1mo agoHugging Face14itsPrerna202 /OctoCodingBench OctoCodingBench: Instruction-Following Benchmark for Coding Agents English | 中文 🌟 Overview OctoCodingBench benchmarks scaffold-aware instruction following in repository-grounded agentic coding. Why OctoCodingBench? Existing benchmarks (SWE-bench, etc.) focus on task completion — whether the agent produces correct code. However, they miss a critical dimension: does the agent follow the rules while solving the task? In real-world agentic coding, agents must… See the full description on the dataset page: https://huggingface.co/datasets/itsPrerna202/OctoCodingBench.texttext-generationn<1K0 likes9 downloads5mo agoHugging Face15itsluketwist /LibHalluBenchgated LibHalluBench - Library Hallucinations Benchmark Welcome to LibHalluBench - the benchmark dataset for testing an LLMs propensity to use non-existent library names during code generation. Using the prompts created in the paper Library Hallucinations in LLM-Generated Code: A Risk Analysis Grounded in Developer Queries, we have curated a dataset of code generation problems that have been observed to trigger a higher rate of hallucinations in LLMs. 📋 dataset | 💾 download | 🤖… See the full description on the dataset page: https://huggingface.co/datasets/itsluketwist/LibHalluBench.texttext-generation1K<n<10K0 likes6 downloads1mo agoHugging Face16itsrocchi /seeweb-llama-it-settexttext-generationn<1K0 likes5 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.