CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ajibawa-2023 /Technical-Architectures-Large Technical Architectures Large (294k Samples) Overview Generating complex, syntactically valid diagram code from natural language requirements is a major challenge for AI models. This dataset bridges that gap by providing over 293,000+ distinct enterprise software architectures generated using two cutting-edge models: GPT-OSS-120B and Qwen3-Coder-Next-FP8. Unlike simple "toy" examples, these architectures model realistic enterprise systems complete with client… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Technical-Architectures-Large.tabulartext-generation100K<n<1M8 likes284 downloads2mo agoHugging Face02techfren /agentbattler-bench AgentBattler Mini Ledger V5 Immutable evidence for 15/15 accepted Mini Ledger V5 runs across 3 harness × model conditions. What is here snapshots/mini-ledger-v5-r5-droid-2026-08-04t13-03-57-476z/site/terminal-campaign.json: compact website and analysis input. snapshots/mini-ledger-v5-r5-droid-2026-08-04t13-03-57-476z/campaign.json: source-revision-preserving campaign index with host paths removed. snapshots/mini-ledger-v5-r5-droid-2026-08-04t13-03-57-476z/runs/:… See the full description on the dataset page: https://huggingface.co/datasets/techfren/agentbattler-bench.tabularn<1K2 likes188 downloads2mo agoHugging Face03Boxoffice1280 /Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques BoxOffice Verified Seeds This dataset contains the released BoxOffice seed datasets used in the benchmark pipeline described in the accompanying paper. The release includes ten verified seeds: 7 11 13 17 19 23 29 31 47 73 For each seed, we provide: a full JSONL file containing warmup rows plus evaluation rows an eval JSONL file containing only the evaluation rows a manifest JSON file a validation JSON file with directional warmup counts Layout viewer/ normalized… See the full description on the dataset page: https://huggingface.co/datasets/Boxoffice1280/Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques.tabulartext-generation1K<n<10K3 likes160 downloads5mo agoHugging Face04BAAI /IndustryInstruction_Technology-Research IndustryInstruction: Technology & Research This repository contains the IndustryInstruction: Technology & Research domain subset of BAAI/IndustryInstruction. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryInstruction: @misc{shi2024industryinstruction, title = {IndustryInstruction}, author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Technology-Research.tabularquestion-answering100K<n<1M0 likes155 downloads1mo agoHugging Face05danielrosehill /Tech-Sentences-For-ASR-Training TechVoice Dataset Work in Progress – This dataset is actively being expanded with new recordings. Dataset Statistics Metric Current Target Progress Duration 38m 43s 5h 0m 0s ██░░░░░░░░░░░░░░░░░░ 12.9% Words 10,412 50,000 ████░░░░░░░░░░░░░░░░ 20.8% Total Recordings: 205 samples Total Characters: 74,312 A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.audioautomatic-speech-recognitionn<1K2 likes142 downloads10mo agoHugging Face06exeyarikus /ru-tech-jobs Russian Tech Jobs Dataset Description Dataset contains job vacancy posts from one of Telegram channels focusing on IT and tech recruitment. Data Fields id: Unique identifier of the post. date: ISO timestamp of when the job was posted. text: The raw text of the post with markdown. views: The view count of the post at the time of scraping. tags: Special tags thats will be taken from text. How to use from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/exeyarikus/ru-tech-jobs.tabulartext-classification1K<n<10K0 likes84 downloads26d agoHugging Face07teplitsa-soc-tech /factbutcher-benchmarkEnglish · Русский FactButcher Russian Fact-Checking Dataset This dataset contains 423 claims in Russian and the results of checking them. Some claims are true, some are false, and some allow more than one defensible answer. You can see what kinds of claims people bring to fact-checkers, investigate a few of them yourself, or use the complete collection to compare different fact-checking tools. A few examples Claim Result Origin Слоны боятся мышей False… See the full description on the dataset page: https://huggingface.co/datasets/teplitsa-soc-tech/factbutcher-benchmark.texttext-classificationn<1K0 likes71 downloads2mo agoHugging Face08buley /breathing-techniques Breathing Techniques 16 evidence-based breathing practices for emotional regulation, with contraindications, session guidance, difficulty levels, and primary benefits. Quick Start from datasets import load_dataset ds = load_dataset("buley/breathing-techniques") print(ds["train"][0]) Structure Field Description id Unique identifier name Technique name category Foundational/Calming, Energizing, Advanced, Specialized difficulty_level Beginner… See the full description on the dataset page: https://huggingface.co/datasets/buley/breathing-techniques.tabulartext-generationn<1K1 likes47 downloads7mo agoHugging Face09guanvireak /khmer-nlp-technical-corpus khmer-nlp-technical-corpus — Khmer Strategic NLP Corpus Dataset Summary This dataset contains peer-grade long-form technical treatises (3,000+ words each) in the Khmer language (km / ភាសាខ្មែរ). Every article is normalized and features neural BiGRU+CRF word segmentation with Zero-Width Space (\u200B) injection to prevent token fragmentation in sub-word tokenizers. Dataset Statistics Total Documents: 3 Train Documents: 3 Total Words: 8,002 Total… See the full description on the dataset page: https://huggingface.co/datasets/guanvireak/khmer-nlp-technical-corpus.tabulartext-generationn<1K0 likes40 downloads9d agoHugging Face10kurdish-tech /KurdishCorpus-clean The Largest Documented Kurmanji and Multi-Dialect Kurdish Dataset — v1.1 A large-scale, multi-dialect Kurdish text corpus prioritizing native Kurmanji (Northern Kurdish) fluency, with substantial Sorani (Central Kurdish) and a Zazaki baseline. Built for language modeling, tokenizer training, and general-purpose Kurdish NLP. This release contains only openly-licensed or presumptively-free redistributable content. A parallel research-tier subset (copyrighted commercial… See the full description on the dataset page: https://huggingface.co/datasets/kurdish-tech/KurdishCorpus-clean.tabulartext-generation1M<n<10M2 likes33 downloads2mo agoHugging Face11Seldon-Technologies /VGIBench VGIBench VGIBench is a video question-answering benchmark of human-validated multiple-choice questions over long-form videos. The questions are designed to mitigate the common mistakes of today's video benchmarks and to expose pragmatic challenges for modern state-of-the-art models. This is the public split (439 questions), released with answers so anyone can score a model with exact-match. A held-out private split is evaluated open-ended by the benchmark maintainers as a… See the full description on the dataset page: https://huggingface.co/datasets/Seldon-Technologies/VGIBench.tabularvisual-question-answeringn<1K0 likes27 downloads2mo agoHugging Face12privacy-tech-lab /ppRegionTesttabularn<1K0 likes20 downloads4y agoHugging Face13privacy-tech-lab /ppAllTesttabularn<1K0 likes19 downloads4y agoHugging Face14TsingyuAI-Tech /ai-detector-ref-entabular10K<n<100K0 likes17 downloads11mo agoHugging Face15hunterbown /bell-labs-technical-archive Bell Labs Documents and Stuff This is a conservative public-release subset of the internal BELLA continued-pretraining corpus. It keeps the Bell-system technical material that survived a stricter final pass for public dataset hosting and removes records that still looked risky, off-scope, or too low-signal for a Hugging Face corpus listing. What is in the release Split Documents train 1220 validation 29 test 42 The release contains 1291 documents out… See the full description on the dataset page: https://huggingface.co/datasets/hunterbown/bell-labs-technical-archive.tabulartext-generation1K<n<10K0 likes15 downloads6mo agoHugging Face16autoshift /Technical-Architectures-Large Technical Architectures Large (210k+ Samples) Overview Generating complex, syntactically valid diagram code from natural language requirements is a major challenge for AI models. This dataset bridges that gap by providing over 210,000 distinct enterprise software architectures generated using two cutting-edge models: GPT-OSS-120B and Qwen3-Coder-Next-FP8. Unlike simple "toy" examples, these architectures model realistic enterprise systems complete with client… See the full description on the dataset page: https://huggingface.co/datasets/autoshift/Technical-Architectures-Large.tabulartext-generation100K<n<1M0 likes15 downloads2mo agoHugging Face17privacy-tech-lab /ppRegionTraintabularn<1K0 likes13 downloads4y agoHugging Face18privacy-tech-lab /cityUnlTraintabular10K<n<100K0 likes12 downloads4y agoHugging Face19privacy-tech-lab /ppLatValtabularn<1K0 likes11 downloads4y agoHugging Face20privacy-tech-lab /ppCityTesttabularn<1K0 likes10 downloads4y agoHugging Face21Phase-Technologies /qubikdatasetsampletabular1K<n<10K0 likes10 downloads3mo agoHugging Face22Dhi-Technologies /amodal-counting-benchmarkgated Amodal Counting Benchmark Product: amodal-counting, "count what detectors can't see": visibility-corrected object counting through crowds, clutter, and occlusion, reported as a calibrated interval rather than a bare point estimate. This dataset is the exact evaluation population the product's own amodal bench command scores against: procedurally generated scenes with known ground-truth occupancy, a simulated detector with a known detectability curve, and the naive-vs-corrected… See the full description on the dataset page: https://huggingface.co/datasets/Dhi-Technologies/amodal-counting-benchmark.tabularobject-detectionn<1K0 likes10 downloads2mo agoHugging Face23privacy-tech-lab /ppCityTraintabularn<1K1 likes9 downloads4y agoHugging Face24privacy-tech-lab /ppLngTraintabularn<1K0 likes9 downloads4y agoHugging Face25privacy-tech-lab /ppAllValtabularn<1K0 likes9 downloads4y agoHugging Face26privacy-tech-lab /ppRegionValtabularn<1K0 likes8 downloads4y agoHugging Face27privacy-tech-lab /ppZipTraintabularn<1K0 likes7 downloads4y agoHugging Face28privacy-tech-lab /ppLatTraintabularn<1K0 likes6 downloads4y agoHugging Face29privacy-tech-lab /ppZipTesttabularn<1K0 likes6 downloads4y agoHugging Face30privacy-tech-lab /ppLatTesttabularn<1K0 likes6 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.