CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01liwu /MNBVCMNBVC: Massive Never-ending BT Vast Chinese corpustext-generation653 likes47k downloads1mo agoHugging Face02wormtooth /MNBVC-judgment这是 MNBVC 项目中裁判文书数据。数据格式使用的是通用数据格式。清洗代码:wormtooth/MNBVC-judgment. tabular1M<n<10M5 likes8.2k downloads3y agoHugging Face03wormtooth /MNBVC-epubs这是 MNBVC 项目中部分 epubs 数据。数据格式使用的是通用数据格式。 data/aliyun-dev: 阿里云开发者社区,数据源:firefox-popkart-org/aliyun-dev data/52pojie: 吾爱破解,数据源:it-ebooks-0/52pojie-2008-2021 data/qdaily: 好奇心日报,数据源:ixinzhi/qdaily-backup data/zuowencom: 作文网,数据源:CyberCommy/zuowencom data/biqu520net: 笔趣网,数据源: biqu520net-10001-20000 biqu520net-20001-30000biqu520net-30001-40000 biqu520net-40001-50000 biqu520net-50001-60000 biqu520net-60001-70000 biqu520net-70001-80000 biqu520net-80001-90000 biqu520net-90001-100000… See the full description on the dataset page: https://huggingface.co/datasets/wormtooth/MNBVC-epubs.tabular1M<n<10M0 likes5k downloads2y agoHugging Face04noeatme /MNBVCMNBVC: Massive Never-ending BT Vast Chinese corpustext-generation0 likes819 downloads7mo agoHugging Face05xianbao /MNBVCMNBVC: Massive Never-ending BT Vast Chinese corpustext-generation0 likes496 downloads4mo agoHugging Face06wanng /wikipedia-zh-mnbvc zhwiki-mnbvc 分项目:爬取并处理中文维基百科语料 数据时间:202302-202305 (持续更新) 主项目:MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集 https://github.com/esbatmop/MNBVC 该项目清洗流程主要参考:https://kexue.fm/archives/4176/comment-page-1 并且使用组员开发的去重工具进行数据格式化。 总行数(样本): 10,754,146 一个示例: { "文件名": "cleaned/zhwiki-20230420/folder_0/723712.txt", "是否待查文件": false, "是否重复文件": false, "文件大小": 558, "simhash": 14363740497821204542, "最长段落长度": 142, "段落数": 6, "去重段落数": 6, "低质量段落数": 0, "段落": [ {… See the full description on the dataset page: https://huggingface.co/datasets/wanng/wikipedia-zh-mnbvc.tabulartext-generation1M<n<10M6 likes381 downloads3y agoHugging Face07semran1 /yulan-code-MNBVC-matlabtext100K<n<1M0 likes350 downloads1y agoHugging Face08gatilin /wenwu_datasets_mnbvctext100K<n<1M0 likes208 downloads1y agoHugging Face09miracleyin /example_mmdata_mnbvc mnbvc mm dataset v2.1 MNBVC 多模态语料数据格式。原链接:https://huggingface.co/datasets/wanng/example_mmdata_mnbvc 参考实现:mm_template_mnbvc 的 mmdata_block.BLOCK_SCHEMA。schema 以那份代码为准,这个数据集是它的示例产物。 字段 字段名称 类型 字段说明 可选 实体ID string 数据的唯一标识符。用于在数据集中确定是哪一条数据。在单个数据集中确定一条数据的实体对象。 必选 md5 string 内容的 md5,用于去重与完整性校验 必选 块ID int32 一个实体对象内的标识符。用于确定一条数据内的一个部分数据。parquet 行的最小单元。 必选 块类型 string 用于保存块的类别。类别的含义为「模态」。取值见下 必选 扩展字段 string 用于保存块的元信息。为可以被成功 load 的 json 字符串。后期可继续扩展 必选… See the full description on the dataset page: https://huggingface.co/datasets/miracleyin/example_mmdata_mnbvc.audioimage-to-textn<1K2 likes64 downloads26d agoHugging Face10qbo-odp /MNBVC-coreMNBVC-core: core split of Massive Never-ending BT Vast Chinese corpus1 likes63 downloads3y agoHugging Face11wanng /example_mmdata_mnbvc这是MNBVC多模态语料小组的图文通用语料的格式展示。 字段说明: 文件md5: 这个字段存储文件的MD5哈希值。MD5是一种广泛使用的哈希函数,它产生一个128位(16字节)的哈希值,通常用于确保数据的完整性。在这里,它可以用来唯一标识文件,或者检查文件是否被更改。 文件id: 文件的唯一标识符。这可以是一个数据库中的主键,或者任何用于唯一标识文件的系统。 页码: 如果数据源是一个多页文档(如PDF文件),这个字段表示文本或图片所在的具体页码。 块id: 在文档中,特定块(block)的唯一标识。这可以用于标识文本或图片所在的具体块(block)。块(block)的定义其实主要用以区分“模态”,所以对于文本来说,可以是多段的文本,也可能是单段的,取决于使用者的实际情况。 文本: 存储文档中的文本内容。这可以是整个文档的文本,或者是特定段落或页面的文本。 图片: 如果文档中包含图片,这个字段存储图片的数据。 时间:… See the full description on the dataset page: https://huggingface.co/datasets/wanng/example_mmdata_mnbvc.tabularn<1K2 likes28 downloads3y agoHugging Face12davidok /mnbvc_Murder_mystery_game0 likes18 downloads2y agoHugging Face13semran1 /yulan-code-MNBVC-googletext1M<n<10M0 likes15 downloads1y agoHugging Face14wangtianxin /MNBVC-QA-with-reporters-from-the-Ministry-of-Foreign-Affairs github 清洗脚本 https://github.com/UnstoppableCurry/MNBVC-QA-with-reporters-from-the-Ministry-of-Foreign-Affair shtml数据清洗 1700个文件,清洗12877条 条外交部记者问数据 清洗前 "<P style="FONT-FAMILY: arial; FONT-SIZE: 14px"  答:当前东亚区域合作总体势头良好,为地区国家抗击疫情和经济复苏提供了积极助力。同时,全球疫情反弹波动,地区热点问题此起彼伏,东亚合作面临更多复杂因素。 /P>" "<P style="FONT-FAMILY: arial; FONT-SIZE: 14px"… See the full description on the dataset page: https://huggingface.co/datasets/wangtianxin/MNBVC-QA-with-reporters-from-the-Ministry-of-Foreign-Affairs.2 likes13 downloads3y agoHugging Face15miracleyin /qiushibaike_mnbvc0 likes5 downloads2y agoHugging Face16zkera /mnbvccx0 likes5 downloads9mo agoHugging Face17longgehaha /mnbvc_ed0 likes4 downloads2y agoHugging Face18mnbvcx /XFUND-LiLThttps://github.com/doc-analysis/XFUND0 likes3 downloads3y agoHugging Face19mnbvcx5643 /Job_Predtext1K<n<10K0 likes2 downloads1y agoHugging Face20quyhhue /mnbvcx0 likes1 downloads9mo agoHugging Face21zkera /mnbvc0 likes1 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.