CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hzb29 /Zhoulifeng-Streaming-Dataset Zhoulifeng-Streaming-Dataset (峰哥直播语料 SFT 数据清洗流水线) 本项目旨在将“峰哥亡命天涯”的原始直播语音转录文本(ASR Transcripts)清洗、提纯、去重并重写,最终构建出高质量的大模型指令微调(SFT)和思维链(CoT)数据集。 本流水线包含 11 个核心处理步骤,通过规则过滤、统计学过滤、语义降噪以及大模型(LLM)深度重写等手段,将原始的五万多条粗糙问答,提纯为高逻辑密度、高“峰味”浓度的高质量训练集。 🗂️ 核心处理流程 (Data Pipeline) 数据清洗呈现严格的“漏斗状”递减,以下是完整的清洗流程序号与功能说明: Step 01: 语音文本标点恢复 (Punctuation Generation) 脚本: 01_punctuation_generate.py 输入/输出: 生成 01_punc_transcripts/ 目录 说明: 针对原始 ASR 语音识别出的无断句纯文本,利用模型恢复正确的标点符号,为后续的语义切分和 QA… See the full description on the dataset page: https://huggingface.co/datasets/hzb29/Zhoulifeng-Streaming-Dataset.textfeature-extraction1K<n<10K7 likes4.1k downloads7mo agoHugging Face02vkatg /streaming-phi-deidentification-benchmark Streaming PHI De-Identification Benchmark Most PHI de-identification benchmarks evaluate a single document in isolation. That is not how clinical data actually moves. A patient's name appears in a clinical note, then in an ASR transcript ten minutes later, then in imaging metadata an hour after that. Each event looks low-risk on its own. The cumulative exposure across modalities is what creates re-identification risk. This dataset captures that. Every record is fully synthetic. It… See the full description on the dataset page: https://huggingface.co/datasets/vkatg/streaming-phi-deidentification-benchmark.documentn<1K0 likes131 downloads7mo agoHugging Face03lijinzhao30 /Streaming_Sampletabularn<1K0 likes47 downloads22d agoHugging Face04electricsheepafrica /africa-synth-streaming-consumption-all African Streaming Consumption | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: json - Sector: technology_digital - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-streaming-consumption-all.tabulartabular-classification10K<n<100K0 likes20 downloads1mo agoHugging Face05Jaiccc /model0_boundary_predict_streaming ## Dataset Card: Terminal Log Boundary Prediction (Streaming) 📋 Dataset Purpose & Model 0 Overview This dataset is designed to train "Model 0" for the Winter 2026 iteration of the AutoDocs project. You can access the official repository here:AutoDocs (Winter 2026) Repository Objective The primary objective of the model is the binary classification of sequential data. It is engineered to process continuous, timestamped terminal logs formatted in XML to… See the full description on the dataset page: https://huggingface.co/datasets/Jaiccc/model0_boundary_predict_streaming.texttext-classification1K<n<10K0 likes12 downloads6mo agoHugging Face06q-future /VQA-stage2-streamingtext1K<n<10K0 likes9 downloads2y agoHugging Face07sourmandaq /yodataset-streaming-elbrustextn<1K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.