CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Hailstone-Technologies /euler-source-parquetstext1M<n<10M0 likes4k downloads3mo agoHugging Face02BAAI /IndustryCorpus_technology[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_technology.texttext-generation10M<n<100M4 likes3.7k downloads1mo agoHugging Face03Hailstone-Technologies /euler-source-parquets-realtext1M<n<10M0 likes3.7k downloads3mo agoHugging Face04QUD-Technologies /quranic-universal-ayahs Qur'anic Universal Ayahs Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset. This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs.audioautomatic-speech-recognition100K<n<1M6 likes2.8k downloads3d agoHugging Face05TechnoBaptist /stack-v3-train 🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/TechnoBaptist/stack-v3-train.tabulartext-generation100M<n<1B0 likes1.2k downloads2mo agoHugging Face06treble-technologies /Treble10-Speech Treble10-Speech (16 kHz) The Treble10-Speech dataset is a dataset for automatic speech recognition (ASR), containing pre-convolved speech files using high fidelity room-acoustic simulations from the Treble10-RIR dataset with 10 different furnished rooms: 2 bathrooms, 2 bedrooms, 2 living rooms with hallway, 2 living rooms without hallway, 2 meeting rooms. The room volumes range between 14 and 46 m3, resulting in reverberation times between 0.17 and 0.84 s. Examples:… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/Treble10-Speech.audioautomatic-speech-recognition1K<n<10K21 likes1.1k downloads11mo agoHugging Face07treble-technologies /Treble10-RIR Treble10-RIR (32 kHz) The Treble10-RIR dataset is a dataset for automatic speech recognition (ASR), containing high fidelity room-acoustic simulations from 10 different furnished rooms: 2 bathrooms, 2 bedrooms, 2 living rooms with hallway, 2 living rooms without hallway, 2 meeting rooms. The room volumes range between 14 and 46 m3, resulting in reverberation times between 0.17 and 0.84 s. Illustrative plots of the rooms and device included in this dataset may be found in the… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/Treble10-RIR.audio1K<n<10K33 likes855 downloads11mo agoHugging Face08Seldon-Technologies /CADBench-Hard CADBench Hard Tasks 43 out of the 105 tasks. Each folder contains the complete task prompt and its authoritative Fusion reference. For all of the tasks, verifiers and sandbox environment, please reach out Dataset categories Domains: computer-aided design, mechanical engineering, and robotics Modalities: natural-language task instructions and native 3D CAD artifacts Use cases: GUI-agent evaluation, computer-use evaluation, reinforcement learning, and deterministic… See the full description on the dataset page: https://huggingface.co/datasets/Seldon-Technologies/CADBench-Hard.textreinforcement-learningn<1K0 likes562 downloads1mo agoHugging Face09Hailstone-Technologies /euler-structural-equationstabular10M<n<100M0 likes434 downloads3mo agoHugging Face10Phase-Technologies /claude-merged-tracestext100K<n<1M0 likes429 downloads2mo agoHugging Face11Technoculture /riddle_senseriddle_sense dataset formatted into an alpaca format dataset for instruction tuning LLMs for reasoning capabilities. textquestion-answering1K<n<10K1 likes413 downloads3y agoHugging Face12gussieIsASuccessfulWarlock /information_technology_instruct_mcq_2481textn<1K2 likes365 downloads2y agoHugging Face13treble-technologies /librispeech_asr_sliced Librispeech Slices Description Librispeech is a large corpus of read English utterances derived from the LibriVox public domain audiobook project. It was assembled to assist in Automatic Speech Recognition tasks, and contains approximately 1000 Hours of utterances recorded at 16kHz. A subset of the original Librispeech dataset was created to support the development of automatic audio scene creation in the Treble SDK environment. To better imitate the natural flow of… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/librispeech_asr_sliced.tabular100K<n<1M0 likes286 downloads7mo agoHugging Face14QUD-Technologies /quran-alignment-benchmark Quran Recitation Alignment Benchmark Audio recordings of Quran recitation with a reviewed word-level ground truth: every recited word, in the order it was recited, with its start and end time, plus the reviewed segmentation and non-Quran regions. This is the corpus behind the Quran Recitation Alignment Benchmark; the task, scoring rules, leaderboard and submission format are documented there, not here. 16 recordings · 357 minutes · 18,421 recited words · Hafs ʿan ʿĀṣim ·… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quran-alignment-benchmark.audioautomatic-speech-recognitionn<1K0 likes274 downloads15d agoHugging Face15Technoculture /chatdoctor-embedded Chat Doctor with Embeddings This dataset is post-processed version of xzuyn/chatdoctor-200k-stripped: Add embeddings for input and output columns using BAAI/bge-small-en-v1.5 Details Sample Count 414k Token Count 1.7b Origin https://drive.google.com/file/d/1lyfqIwlLSClhgrCutWuEe_IACNq6XNUt/view Source of raw data ? Processing details paper Embedding Model BAAI/bge-small-en-v1.5 Data Diversity index Example Output GPT-4 Rationale GPT-4… See the full description on the dataset page: https://huggingface.co/datasets/Technoculture/chatdoctor-embedded.text100K<n<1M3 likes248 downloads3y agoHugging Face16electricsheepasia /asia-science-technology-world-bank-science-and-technology-indica Maldives - Science and Technology Publisher: World Bank Group · Source: HDX · License: cc-by · Updated: 2026-04-28 Abstract Contains data from the World Bank's data portal. There is also a consolidated country dataset on HDX. Technological innovation, often fueled by governments, drives industrial growth and helps raise living standards. Data here aims to shed light on countries technology base: research and development, scientific and technical journal articles… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-science-technology-world-bank-science-and-technology-indica.tabulartabular-classificationn<1K0 likes244 downloads5mo agoHugging Face17Hailstone-Technologies /harmonia-infinity-corpustext100M<n<1B0 likes228 downloads5mo agoHugging Face18Hailstone-Technologies /harmonia-triples-rust-code-traversal harmonia-triples-rust Triples for source rust emitted by the ingest pipeline (current wave: v0.7). Schema: (s, p, o, src) with full provenance per ADR-0011. Pre-HHEC. Provenance Each parquet shard carries the full provenance chain per ADR-0011: s, p, o, src columns (when this is a triples-stage dataset) src = "<dataset>:<version>:<file>" for triples Causal registry events recorded at causal_registry/master.jsonl chain Architecture Part of Harmonia… See the full description on the dataset page: https://huggingface.co/datasets/Hailstone-Technologies/harmonia-triples-rust-code-traversal.textgraph-ml1M<n<10M0 likes199 downloads5mo agoHugging Face19QTE-Technologies /industrial-technical-archive 🚀 Latest Updates (July, 2026) Version: v07.2026 (Verified) Status: Integrated with 1,000,000+ records. New Files: product-E-26-07-2026.csv & product-V-26-07-2026.csv. QTE Technologies: Industrial & Scientific Knowledge Base Wikidata Entity: Q138411149 IPFS CID: bafybeibogxxuhmzfrsuhcfd4qr4tmc4okhmrcwhp3266hq47ccuyjnjxoq Official Neural Hub: qtetech.github.io This is the permanent technical archive for QTE Technologies, ensuring long-term accessibility of… See the full description on the dataset page: https://huggingface.co/datasets/QTE-Technologies/industrial-technical-archive.image10K<n<100K0 likes196 downloads2mo agoHugging Face20electricsheepafrica /africa-owid-access-to-clean-fuels-and-technologies-for-cooking Access To Clean Fuels And Technologies For Cooking | Africa (Our World in Data) | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: parquet - Sector: other_unclassified - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-owid-access-to-clean-fuels-and-technologies-for-cooking.tabulartabular-classification1K<n<10K0 likes189 downloads1mo agoHugging Face21BAAI /IndustryInstruction_Technology-Research IndustryInstruction: Technology & Research This repository contains the IndustryInstruction: Technology & Research domain subset of BAAI/IndustryInstruction. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryInstruction: @misc{shi2024industryinstruction, title = {IndustryInstruction}, author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Technology-Research.tabularquestion-answering100K<n<1M0 likes155 downloads1mo agoHugging Face22Hailstone-Technologies /harmonia-triples-stackexchange-document-traversal harmonia-triples-stackexchange-slice Triples for source stackexchange-slice emitted by the ingest pipeline (current wave: v0.6). Schema: (s, p, o, src) with full provenance per ADR-0011. Pre-HHEC. Provenance Each parquet shard carries the full provenance chain per ADR-0011: s, p, o, src columns (when this is a triples-stage dataset) src = "<dataset>:<version>:<file>" for triples Causal registry events recorded at causal_registry/master.jsonl chain… See the full description on the dataset page: https://huggingface.co/datasets/Hailstone-Technologies/harmonia-triples-stackexchange-document-traversal.textgraph-ml100K<n<1M0 likes138 downloads5mo agoHugging Face23Phase-Technologies /forge-3b-dpo-data FORGE-3B DPO Preference Data Tokenized (prompt, chosen, rejected) preference triples for DPO post-training of FORGE-3B, built per the FORGE paper Section 6.2 / Appendix A.2. This is data preparation output only — no model was trained to produce this. Stats Total pairs: 0 (paper target: ~200,000) Domains: 0/4 Context length: 4096 tokens (paper Appendix A.2, DPO block) Format: unpacked — one (prompt, chosen, rejected) triple per training example Chat template:… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-dpo-data.texttext-generation100K<n<1M0 likes138 downloads3mo agoHugging Face24system-technologies /MedCase-Structured MedCase-Structured Dataset for Paper MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settings Structured FHIR R4 representations of clinical reasoning cases, derived from the MedCaseReasoning dataset (Wu et al., 2025). Each case pairs a free-text clinical presentation with a machine-readable FHIR bundle and a held-out ground-truth diagnosis, supporting evaluation of clinical information extraction, terminology coding… See the full description on the dataset page: https://huggingface.co/datasets/system-technologies/MedCase-Structured.text1K<n<10K1 likes132 downloads3mo agoHugging Face25Hailstone-Technologies /harmonia-graph-causal harmonia-graph-causal Causal graph triples from structured sources: bnlearn Bayesian networks, Reactome pathways, STRING protein interactions, SIGNOR signaling, Wikidata causal properties, ConceptNet causal relations, Tübingen cause-effect pairs, and 60+ Harmonia Structural Causal Models covering the full human experience. Schema Every row: s (subject), p (predicate), o (object), src (provenance) src format: dataset:version:file Part of the Harmonia… See the full description on the dataset page: https://huggingface.co/datasets/Hailstone-Technologies/harmonia-graph-causal.textgraph-ml1M<n<10M0 likes122 downloads5mo agoHugging Face26Technoculture /synthetic-clinical-notes-embedded Synthetic Clinical Notes This dataset is post-processed version of starmpcc/Asclepius-Synthetic-Clinical-Notes: Turn into Alpaca format (instruction, input, and output) Add embeddings for input and output columns using BAAI/bge-small-en-v1.5 Details Sample Count 158k Token Count 648m Origin https://figshare.com/authors/Zhengyun_Zhao/16480335 Source of raw data PubMed Central (PMC) and MIMIC 3 Processing details original, paper Embedding Model… See the full description on the dataset page: https://huggingface.co/datasets/Technoculture/synthetic-clinical-notes-embedded.textquestion-answering100K<n<1M10 likes88 downloads3y agoHugging Face27rickphd /gold-reddit-ai-technology-corpus Gold Reddit AI and Technology Corpus Dataset Authors Ricardo Flores-Moyano, Universidad San Francisco de Quito Felipe Rosero-Polo, Universidad San Francisco de Quito José Vega-Sánchez, Universidad San Francisco de Quito Maria Baldeon-Calisto, Wake Forest University Dataset Summary This dataset provides a curated Gold corpus of 1,614 Reddit posts that were publicly accessible at collection time. The corpus is primarily English, with a smaller… See the full description on the dataset page: https://huggingface.co/datasets/rickphd/gold-reddit-ai-technology-corpus.tabulartext-classification1K<n<10K0 likes85 downloads23h agoHugging Face28bunkalab /medium-sample-technologySample with the keyword "Technology" taken from https://huggingface.co/datasets/fabiochiu/medium-articles text1K<n<10K0 likes76 downloads3y agoHugging Face29Technoculture /medical-prescriptionsimagen<1K7 likes68 downloads2y agoHugging Face30WallisF /Lab-1-language-technologytext10K<n<100K0 likes68 downloads14d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.