CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fewshot-goes-multilingual /cs_csfd-movie-reviews Dataset Card for CSFD movie reviews (Czech) Dataset Description The dataset contains user reviews from Czech/Slovak movie databse website https://csfd.cz. Each review contains text, rating, date, and basic information about the movie (or TV series). The dataset has in total (train+validation+test) 30,000 reviews. The data is balanced - each rating has approximately the same frequency. Dataset Features Each sample contains: review_id: unique string identifier… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_csfd-movie-reviews.texttext-classification10K<n<100K2 likes373 downloads4y agoHugging Face02DianJin /DianJin-CSC-Data Qwen DianJin Platform | Github | ModelScope | Paper 📢 Introduction Effective customer support requires not only accurate problem-solving but also structured and empathetic communication aligned with professional standards. However, existing dialogue datasets often lack strategic guidance, and realworld service data is difficult to access and annotate. To address this, we introduce the task of Customer Support Conversation (CSC)… See the full description on the dataset page: https://huggingface.co/datasets/DianJin/DianJin-CSC-Data.text10K<n<100K8 likes304 downloads1y agoHugging Face03shibing624 /CSC Dataset Card for CSC 中文拼写纠错数据集 Repository: https://github.com/shibing624/pycorrector Dataset Description Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts. CSC is challenging since many Chinese characters are visually or phonologically similar but with quite different semantic meanings. 中文拼写纠错数据集,共27万条,是通过原始SIGHAN13、14、15年数据集和Wang271k数据集合并整理后得到,json格式,带错误字符位置信息。 Original Dataset Summary test.json 和… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/CSC.texttext-generation100K<n<1M37 likes218 downloads3y agoHugging Face04fewshot-goes-multilingual /cs_czech-named-entity-corpus_2.0 Dataset Card for Czech Named Entity Corpus 2.0 Dataset Description The dataset contains Czech sentences and annotated named entities. Total number of sentences is around 9,000 and total number of entities is around 34,000. (Total means train + validation + test) Dataset Features Each sample contains: text: source sentence entities: list of selected entities. Each entity contains: category_id: string identifier of the entity category category_str:… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_czech-named-entity-corpus_2.0.texttoken-classification1K<n<10K3 likes169 downloads4y agoHugging Face05Macropodus /csc_clean_wang271k csc_eval_public 一、测评数据说明 1.1 数据清洗 余-馀: 替换为馀-余 other - 馀: 替换为余 覆-复: 替换为复-覆 other-覆: # 答疆/回覆/反覆 # 覆审 他-她:不纠 她-他:不纠 人名不纠: 识别人名并丢弃 的得地: 建议丢弃(标注得不准) # # 的 - 地 # # 的 - 得 # # 它 - 他 # # 哪 - 那 # # 改-大小改: 余-馀 覆-复 借-藉 功-工 琅-瑯 震-振 百-白 也-叶 经-禁(经不起-禁不起) # # 部分不变(人名): 小-晓 一-逸 佳-家 得-地(马哈得) 红-虹 民-明 # # 匹配上但是不改的: 惟-唯 象-像 查-察 立-利 止-只 建-健 他-它 地-的 定-订 带-戴 力-利 成-城 点-店 # # 匹配上但是不改的: 作-做 得-的 场-厂 身-生 有-由 种-重 理-里 # # 空白没匹配上: 今-在 年-今 前-目 当-在 目-在 者-是 # # 外国人名等:其-齐 课-科 博-波… See the full description on the dataset page: https://huggingface.co/datasets/Macropodus/csc_clean_wang271k.text100K<n<1M2 likes84 downloads2y agoHugging Face06csc-architecture /csc-wireless-latency-synthetic-100k CSC Wireless Latency Synthetic Dataset (100k) This synthetic dataset provides 100,000 prompt-completion pairs designed for training and evaluating PHY/MAC cross-layer optimization models in hybrid Li-Fi/RF wireless networks. Official Core Implementation & Runtime To parse, simulate, or process this dataset according to the official protocol specifications, please utilize the official runtime library: Core Protocol Library (npm):… See the full description on the dataset page: https://huggingface.co/datasets/csc-architecture/csc-wireless-latency-synthetic-100k.texttext-generation100K<n<1M0 likes56 downloads7d agoHugging Face07shibing624 /CSC-gpt4 Dataset Card for Chinese Spelling Correction(gpt4 fixed version) 中文拼写纠错数据集 Repository: https://github.com/shibing624/pycorrector Dataset Description Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts. CSC is challenging since many Chinese characters are visually or phonologically similar but with quite different semantic meanings.… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/CSC-gpt4.texttext-generation1K<n<10K4 likes48 downloads2y agoHugging Face08csc-architecture /csc-decision-intelligence-dataset CSC Decision Intelligence Dataset Deterministic decision intelligence seeds and cryptographic verification samples for multi-dimensional evaluation protocols. Dataset Description This dataset provides deterministic baseline seeds used by the CSC Protocol (@csc-protocol/core) to evaluate institutional, corporate, and healthcare entities under autonomous AI governance rules. Supported Domains Healthcare (medical): Facility operational efficiency… See the full description on the dataset page: https://huggingface.co/datasets/csc-architecture/csc-decision-intelligence-dataset.texttabular-classificationn<1K0 likes40 downloads14d agoHugging Face09fewshot-goes-multilingual /cs_czech-court-decisions-ner Dataset Card for Czech Court Decisions NER Dataset Description Czech Court Decisions NER is a dataset of 300 court decisions published by The Supreme Court of the Czech Republic and the Constitutional Court of the Czech Republic. In the documents, 4 types of named entities are selected. Dataset Features Each sample contains: filename: file name in the original dataset text: court decision document in plain text entities: list of selected entities. Each entity… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_czech-court-decisions-ner.texttoken-classificationn<1K2 likes38 downloads4y agoHugging Face10astromindinc /axia-csc-corpus Axia — Chandra Source Catalog corpus astromindinc/axia-csc-corpus — 51,450 X-ray sources from the Chandra Source Catalog 2.1, each carrying: Per-photon event lists in two forms: an event_list pruned to a single 8 h window in 0.5-8 keV (the input shape the Axia fine-tuned model trained on), and an original_event_list containing the full unpruned observation (the input to the model-free spectrum-snapshot / light-curve pipeline). A 64-d learned embedding (pca_64d) suitable for… See the full description on the dataset page: https://huggingface.co/datasets/astromindinc/axia-csc-corpus.tabular10K<n<100K0 likes28 downloads21d agoHugging Face11henryy1990 /CSC Dataset Card for CSC 中文拼写纠错数据集 Repository: https://github.com/shibing624/pycorrector Dataset Description Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts. CSC is challenging since many Chinese characters are visually or phonologically similar but with quite different semantic meanings. 中文拼写纠错数据集,共27万条,是通过原始SIGHAN13、14、15年数据集和Wang271k数据集合并整理后得到,json格式,带错误字符位置信息。 Original Dataset Summary test.json 和… See the full description on the dataset page: https://huggingface.co/datasets/henryy1990/CSC.texttext-generation100K<n<1M0 likes23 downloads10mo agoHugging Face12Macropodus /csc_public_de3 csc_public_de3数据集 数据来源 1.由人民日报/学习强国/chinese-poetry等高质量数据人工生成; 2.来自人民日报高质量语料; 3.来自学习强国网站的高质量语料; 4.源数据为qwen生成的好词好句; 5.古诗词chinese-poetry; 文言文garychowcmu/daizhigev20; 数据简介 该数据主要为'的地得'纠错; 其中训练数据130753条, 验证数据5545条, 测试数据5545条; 句子平均长度为36, 最长句子长度为414, 最短为5, 95%的为89, 75%的为46, 60%的为34; 每个句子中字的平均错误数为2; 数据详情 ################################################################################################################################ train.json 130753… See the full description on the dataset page: https://huggingface.co/datasets/Macropodus/csc_public_de3.texttext-generation100K<n<1M1 likes22 downloads2y agoHugging Face13RuiWu1123Alignment /CS-Chain-4ktext1K<n<10K0 likes2 downloads1y agoHugging Face14CCSSE1 /cscstextn<1K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.