CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fatihdx /tr-rss-haber-akisi-verisi TR-RSS Haber Akışı Verisi TL;DR — Bu veri seti, Türkiye odaklı haber/RSS akışlarından toplanan kayıtları; mükerrerlik, spam, reklam, amaç dışı kategori, yurtdışı odak ve editoryal çerçeve yoğunluğu açısından katmanlı kalite kontrolden geçirerek erken sinyal üretimine uygun hâle getirir. Doğrulama kararı / verdict üretmez; ClaimReview ve dezenformasyon araştırmaları için upstream izleme ve kaynak önceliklendirme katmanı olarak tasarlanmıştır. Ölçek: 307.800 öğe incelendi →… See the full description on the dataset page: https://huggingface.co/datasets/fatihdx/tr-rss-haber-akisi-verisi.tabulartext-classification10K<n<100K0 likes204 downloads3mo agoHugging Face02fatcat55 /delvantic-stock-knowledge-layer Delvantic Stock Knowledge Layer A 872k-word, source-cited textbook of stock analysis and trading, organized as a tree — the reference layer behind a live AI research engine, published in full. Every finance dataset on the Hub is numbers: prices, filings, labelled headlines. This is the missing other half — the explanations. 771 documents on how the machinery of markets actually works, from reading a cash-flow statement to why volatility regimes break strategies, each one written… See the full description on the dataset page: https://huggingface.co/datasets/fatcat55/delvantic-stock-knowledge-layer.tabulartext-retrieval1K<n<10K0 likes141 downloads1mo agoHugging Face03Fatema142 /BB-FinQA-X BB-FinQA-X BB-FinQA-X is a 500-item, expert-grounded question-answering benchmark built from the Bangladesh Bank Annual Report, FY2024–25, the central bank of Bangladesh's official yearly report on macroeconomic conditions, monetary policy, banking-sector supervision, and financial markets. Every question is paired with a literal, page-cited evidence quote from the source report, drawn from narrative text, statistical tables, and charts alike. This makes the dataset suitable for… See the full description on the dataset page: https://huggingface.co/datasets/Fatema142/BB-FinQA-X.texttable-question-answeringn<1K0 likes98 downloads19d agoHugging Face04fatty-belly /MinecraftSkillDiscoveryThis is the segmented datasets of the project presented in the paper Open-World Skill Discovery from Unsegmented Demonstrations. Code: https://github.com/CraftJarvis/SkillDiscovery Project Page: https://craftjarvis.github.io/SkillDiscovery Each line of the jsonl file consists of the video file name and the boundaries [begin1, end1], [begin2, end2], ... Events information is also included in the "with info" file. The video files can be downloaded here. Notice that we use the 7.x version. textrobotics10K<n<100K0 likes71 downloads1y agoHugging Face05sermonindex /early-church-fathers Early Church Fathers — Scripture Citation Index 68,240 passages from 349 Church Fathers, each keyed to the Bible verse it comments on. Drawn from 20,253 distinct works and covering all 66 books. This is a patristic catena in machine-readable form: given a verse, it returns what the Fathers said about it. Nothing comparable exists as an open dataset — the underlying translations are freely available, but the verse-level alignment is the work, and that is what this releases.… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/early-church-fathers.tabulartext-retrieval10K<n<100K0 likes69 downloads16d agoHugging Face06fatihdx /tr-rss-haber-akisi-schema TR-RSS Haber Akışı Şeması TR-RSS Haber Akışı Şeması, Android/Termux üzerinde çalışan RSS-to-Telegram haber akışı prototipinde kullanılmak üzere hazırlanmış örnek kaynak, anahtar kelime, kategori ve geri bildirim veri şemalarını içerir. Bu çalışma; Türkçe RSS kaynaklarından gelen haber başlıklarının kural tabanlı filtreleme, kategori eşleştirme ve kullanıcı geri bildirimiyle daha işlevsel bir bilgi akışına dönüştürülmesini amaçlayan bağımsız bir açık veri/dokümantasyon… See the full description on the dataset page: https://huggingface.co/datasets/fatihdx/tr-rss-haber-akisi-schema.texttext-classificationn<1K0 likes56 downloads3mo agoHugging Face07Fatema142 /LabourActQA LabourActQA LabourActQA is a 500-item, expert-verified question-answering benchmark built directly from the statutory text of the Bangladesh Labour Act, 2006 (as amended). Every question, reference answer, and supporting evidence quote is drafted and cross-checked against the Act's own sections, subsections, provisos, and cross-references; no case law, commentary, or secondary legal literature is used at any stage. The dataset spans seven reasoning categories across three… See the full description on the dataset page: https://huggingface.co/datasets/Fatema142/LabourActQA.textquestion-answeringn<1K0 likes44 downloads13d agoHugging Face08FractalAIResearch /Fathom-V0.4-RL-Compressiontexttext-generation1K<n<10K1 likes40 downloads1y agoHugging Face09Mawube /fatima-audio-perturbations Audio Perturbation TTS Gold Sentences A curated set of English sentences for perturbation-based blind-spot evaluation of audio-LLM judges on synthesised speech. Overview This dataset provides original (clean) sentences sampled from established TTS benchmarks. The sentences are designed to be fed through TTS models to generate clean audio (A_gold), then perturbed at the text level (S_gold → S_pert) and re-synthesised (A_pert) to test whether audio-LLM judges can… See the full description on the dataset page: https://huggingface.co/datasets/Mawube/fatima-audio-perturbations.texttext-to-speech1K<n<10K0 likes34 downloads1mo agoHugging Face10fatdogLeader /promql-poc-1textn<1K1 likes30 downloads2y agoHugging Face11fathyshalab /google-prestotext100K<n<1M2 likes29 downloads4y agoHugging Face12fathurfrs /qna-hukum-indonesiatext1K<n<10K1 likes28 downloads2y agoHugging Face13MYGBM /fatima-fellowship-blindspots Dataset Card: Tiny-Aya-Base Amharic Evaluation Blindspots Dataset Description This dataset provides a targeted, interpretable checklist of reasoning failures and blindspots discovered in the CohereLabs/tiny-aya-base model when evaluated on Amharic language tasks across arithmetic, logic, science, history, and geography domains. As models scale, evaluating their cross-lingual reasoning capabilities requires moving beyond aggregate metrics. This repository adopts a… See the full description on the dataset page: https://huggingface.co/datasets/MYGBM/fatima-fellowship-blindspots.textn<1K0 likes28 downloads7mo agoHugging Face14FatimaAfzal01 /smollm3-3b-base-blind-spots SmolLM3-3B-Base Blind Spots Dataset This dataset contains 10 test cases where I explored the failure modes of SmolLM3-3B-Base, a 3 billion parameter base language model released by HuggingFace in 2025. The goal was to find diverse cases where the model makes clearly incorrect or unexpected completions its "blind spots." Model Tested Model: HuggingFaceTB/SmolLM3-3B-Base Parameters: 3B Type: Base pretrained model License: Apache 2.0 How I Loaded the Model I… See the full description on the dataset page: https://huggingface.co/datasets/FatimaAfzal01/smollm3-3b-base-blind-spots.texttext-generationn<1K0 likes23 downloads7mo agoHugging Face15FateDefier /MineSafety-QA-Dataset 矿山安全领域 QA 数据集 基于中国矿山安全法规构建的问答对数据集,用于 QLoRA 领域微调。 数据来源 《煤矿安全规程》(2025) 《金属非金属矿山安全规程》(2020) 数据规模 原始生成:7874 条 AI 质量评估过滤后:7265 条 数据格式 Alpaca 格式,包含 <think> 推理链: { "instruction": "问题", "input": "", "output": "<think>\n推理过程...\n</think>\n\n正式回答...", "system": "你是一位精通中国矿山安全法律法规的资深专家..." } 构建流程 PDF 规程文档经 MinerU 转为 Markdown Easy Dataset 自动分块、提取问题、生成答案(DeepSeek-R1-0528-Qwen3-8B) AI 自动评分(满分 5 分),过滤 3.5 分以下的低质量 QA 对… See the full description on the dataset page: https://huggingface.co/datasets/FateDefier/MineSafety-QA-Dataset.textquestion-answering1K<n<10K0 likes21 downloads4mo agoHugging Face16mostafaamiri /fa-topic-sentences README for fa-topic-sentences Dataset Overview The fa-topic-sentences dataset is a comprehensive collection of sentences categorized into various topics. Each topic contains approximately 50 sentences in Persian, accompanied by a paraphrased version of each sentence. The dataset is structured in JSON format, providing a straightforward method for accessing individual entries. Topics Included The dataset encompasses the following topics: History Fashion… See the full description on the dataset page: https://huggingface.co/datasets/mostafaamiri/fa-topic-sentences.textn<1K0 likes19 downloads2y agoHugging Face17fatzin1 /copytextn<1K0 likes14 downloads1y agoHugging Face18Nabeelah04 /fatima_institute_blind_spot Blind Spots of Nanbeige/Nanbeige4-3B-Base 1. Model Tested Nanbeige/Nanbeige4-3B-Base Field Detail Released December 13, 2025 Parameters ~3B Type TRUE BASE MODEL — pre-trained only on 23 trillion tokens, no SFT, no RLHF Languages English + Chinese (primary), multilingual coverage License Apache 2.0 2. How the Model Was Loaded The model was loaded on Google Colab (free tier, T4 GPU, 16 GB VRAM) using the Hugging Face transformers… See the full description on the dataset page: https://huggingface.co/datasets/Nabeelah04/fatima_institute_blind_spot.texttext-generationn<1K0 likes14 downloads7mo agoHugging Face19fatihburakkaragoz /evliya-celebi-seyahatname-ocr Evliya Celebi Seyahatname OCR Corpus OCR-derived text from seven volumes of Evliya Celebi's Seyahatname, packaged for corpus exploration, language modeling, OCR-quality analysis, and historical Ottoman Turkish / Turkish NLP work. Configs pages: one row per OCR page, with page numbers and OCR status. documents: one row per available volume, with page text concatenated. Coverage Available books: 1, 3, 4, 6, 7, 9, 10. Missing from the 1-10 sequence: 2, 5, 8.… See the full description on the dataset page: https://huggingface.co/datasets/fatihburakkaragoz/evliya-celebi-seyahatname-ocr.tabulartext-generation1K<n<10K1 likes14 downloads5mo agoHugging Face20fatdogLeader /promQL-traintextn<1K5 likes13 downloads2y agoHugging Face21Muhammad-Ikhwan-Fathulloh /Saka-Eval Saka-Eval Benchmark Dataset Dataset Description Saka-Eval is a curated benchmark dataset for evaluating Indonesian NLP models, particularly focused on colloquial language, public services, and government entities. This dataset was generated as part of the Saka-NLP ecosystem to provide a standardized way to measure model performance in real-world Indonesian contexts. Task Summaries Sentiment Analysis (sentiment): 100 samples of public service… See the full description on the dataset page: https://huggingface.co/datasets/Muhammad-Ikhwan-Fathulloh/Saka-Eval.texttext-classificationn<1K0 likes13 downloads4mo agoHugging Face22faton1995 /cad_datasettextn<1K0 likes10 downloads2y agoHugging Face23abdulmatinomotoso /Fatimah_Fellowship_Blind_Spot Qwen3-4B-Base Blind Spots Dataset Model Tested Model: Qwen/Qwen3-4B-Base Type: Causal Language Model — pretrained base model (NOT instruction-tuned) Parameters: 4.0 billion (3.6B non-embedding) Architecture: 36 layers, 32 attention heads (GQA: 32 Q / 8 KV) Context Length: 32,768 tokens Training: 36 trillion tokens across 119 languages in a 3-stage pretraining pipeline Overview This dataset documents 10 confirmed blind spots of Qwen3-4B-Base identified… See the full description on the dataset page: https://huggingface.co/datasets/abdulmatinomotoso/Fatimah_Fellowship_Blind_Spot.texttext-generationn<1K0 likes10 downloads7mo agoHugging Face24Amman-shah /WooCommerce-Fatal-Error-Core-Patch-Datasetgated WooCommerce Fatal Error & Core Patch Dataset Auto-generated training dataset for WordPress/WooCommerce error resolution. Stats Total samples: 126 Generated by: NexusOS v5.0 Niche: wordpress_woocommerce Format Each sample contains: instruction: The error or problem description output: The solution/fix grade: Quality grade (A/B/C) score: Quality score (0-10) textn<1K0 likes7 downloads4mo agoHugging Face25Fathi773 /Priestess_Arknightstextn<1K0 likes6 downloads5mo agoHugging Face26FatenAbdullatif /insurance-charge-mlops-logstabularn<1K0 likes4 downloads2y agoHugging Face27FatumMaster /test_GORtextn<1K0 likes4 downloads2y agoHugging Face28fatmahamad /Chatgptttextn<1K0 likes3 downloads3y agoHugging Face29jasonlinzi /fatymttest1textn<1K0 likes3 downloads2y agoHugging Face30fatmerajabi11 /alignment_resultstextn<1K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.