CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01taiwan-corpora /twngrams Taiwanese Mandarin web n-grams Word 1–4-gram counts for Taiwanese Mandarin* (臺灣華語, cmn-Hant-TW), computed over the Taiwan slice of a large web crawl after variety filtering by twfilter 0.1.0 with the published twfilter-tables: every sentence behind these counts passed the 教育部 character-inventory gate, the simplified-character round-trip, the mainland-orthography, mainland-lexicon, written-Cantonese, Hong Kong and Singapore detectors, and block-level evidence of Taiwan-specific… See the full description on the dataset page: https://huggingface.co/datasets/taiwan-corpora/twngrams.texttext-generation1M<n<10M0 likes158 downloads28d agoHugging Face02Moo /korean-parallel-corporatexttranslation10K<n<100K21 likes136 downloads4y agoHugging Face03SaiedAlshahrani /Wikipedia-Corpora-Report Dataset Card for "Wikipedia-Corpora-Report" This dataset is used as a metadata database for the online WIKIPEDIA CORPORA META REPORT dashboard that illustrates how humans and bots generate or edit Wikipedia editions and provides metrics for “pages” and “edits” for all Wikipedia editions (320 languages). The “pages” metric counts articles and non-articles, while the “edits” metric tallies edits on articles and non-articles, all categorized by contributor type: humans or bots. The… See the full description on the dataset page: https://huggingface.co/datasets/SaiedAlshahrani/Wikipedia-Corpora-Report.text1K<n<10K0 likes68 downloads3y agoHugging Face04sello-ralethe /SA-Parallel-Corpora SA-Parallel-Corpora Sentence-aligned English to isiZulu, isiXhosa, Sesotho and Sepedi bitext, drawn from South African government publications. Produced for the doctoral thesis Injecting Commonsense Knowledge into Pretrained Language Models for Low Resource Languages (University of Cape Town, 2026). Code at https://github.com/sello-ralethe/SA-knowledge Structure One configuration per language pair, each with train, validation and test splits. Splits are assigned… See the full description on the dataset page: https://huggingface.co/datasets/sello-ralethe/SA-Parallel-Corpora.tabular10K<n<100K0 likes68 downloads14d agoHugging Face05line-corporation /JIC-VQA JIC-VQA Dataset Description Japanese Image Classification Visual Question Answering (JIC-VQA) is a benchmark for evaluating Japanese Vision-Language Models (VLMs). We built this benchmark based on the recruit-jp/japanese-image-classification-evaluation-dataset by adding questions to each sample. All questions are multiple-choice, each with four options. We select options that closely relate to their respective labels in order to increase the task's difficulty. The… See the full description on the dataset page: https://huggingface.co/datasets/line-corporation/JIC-VQA.imagevisual-question-answering1K<n<10K4 likes66 downloads2y agoHugging Face06QIRIM /crh-parallel-corpora-document-level-noisytabulartranslation10K<n<100K1 likes52 downloads2y agoHugging Face07Horeknad /komi-russian-parallel-corpora Source Datasets 1 - news from the website of the Komi administration (https://rkomi.ru/) 2 - Komi media library (http://videocorpora.ru/) 3 - Millet porridge by Ivan Toropov (adaptation) Authors Shilova Nadezhda Chernousov Georgy texttranslation10K<n<100K2 likes51 downloads3y agoHugging Face08corpstacking /corporate-crypto-treasuries Corporate Crypto Treasury Holdings Every company, government and ETF known to hold Bitcoin, Ethereum or Solana on its balance sheet, with the primary source document behind each figure. 359 positions across 297 entities in 39 countries. 358 of the 359 carry a link to the filing or disclosure the number came from. Maintained by CorpStacking. What makes this different from a price feed Most crypto datasets are market data. This one is balance-sheet data read out of… See the full description on the dataset page: https://huggingface.co/datasets/corpstacking/corporate-crypto-treasuries.tabulartabular-regressionn<1K0 likes48 downloads5d agoHugging Face09QIRIM /crh-parallel-corporagated Crimean Tatar-English/Russian/Ukrainian Parallel Corpora Overview This repository contains the Crimean Tatar-English/Russian/Ukrainian parallel corpora, a collection of sentences in Crimean Tatar and their corresponding translations in English/Russian/Ukrainian. The dataset is intended for use in natural language processing (NLP) tasks such as machine translation and cross-lingual analysis. All sentences in Crimean Tatar are written in Cyrillic and/or Latin script. If any… See the full description on the dataset page: https://huggingface.co/datasets/QIRIM/crh-parallel-corpora.texttranslation100K<n<1M2 likes16 downloads2y agoHugging Face10Arseniy-Sandalov /Georgian-Parallel-Corpora English-Georgian Parallel Dataset 🇬🇧🇬🇪 📄 Dataset Overview The English-Georgian Parallel Dataset is sourced from OPUS, a widely used open collection of parallel corpora. This dataset contains aligned sentence pairs in English and Georgian, extracted from Wikipedia translations. Corpus Name: Wikimedia Package: wikimedia.en-ka (Moses format) Publisher: OPUS (Open Parallel Corpus) Release: v20230407 Release Date: April 13, 2023 License: CC–BY-SA 4.0 🔗… See the full description on the dataset page: https://huggingface.co/datasets/Arseniy-Sandalov/Georgian-Parallel-Corpora.texttranslation10K<n<100K0 likes8 downloads2y agoHugging Face11Shoriful025 /corporate_esg_risk_analyticstabularn<1K0 likes8 downloads9mo agoHugging Face12Siyam025 /Corporate_Employee_Profilestabularn<1K0 likes5 downloads9mo agoHugging Face13St4n /en_corpora_parliament_processedtext1M<n<10M0 likes4 downloads3y agoHugging Face14jason1966 /agewerc_corporate-credit-rating Corporate Credit Rating Credit Ratings of Big US Firms and their Financials Dataset Info Source: Kaggle Original Size: 0.29 MB Kaggle Downloads: 4,884 Files: 1 Files corporate_rating.csv Mirrored from Kaggle tabular1K<n<10K0 likes4 downloads6mo agoHugging Face15upb-nlp /metra_western_european_drama_annotated_corporatext1M<n<10M0 likes4 downloads3mo agoHugging Face16infinite-dataset-hub /CorporateMailCategorization CorporateMailCategorization tags: mails, classification, business Note: This is an AI-generated dataset so its content may be inaccurate or false Dataset Description: The 'CorporateMailCategorization' dataset is a collection of CEO office letters and business-related mails. Each entry in the dataset includes the original email text and a label that classifies the email's content into one of several categories relevant to corporate communication. This dataset can be used to train… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/CorporateMailCategorization.textn<1K0 likes3 downloads2y agoHugging Face17infinite-dataset-hub /CorporateCommMail CorporateCommMail tags: Email Classification, Priority, BusinessCommunication Note: This is an AI-generated dataset so its content may be inaccurate or false Dataset Description: The 'CorporateCommMail' dataset is a curated collection of business communication emails with a focus on understanding the priority levels assigned to emails based on employee responses. The dataset includes various features such as sender, subject, body, timestamp, receiver, CC (carbon copy recipients)… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/CorporateCommMail.textn<1K0 likes3 downloads2y agoHugging Face18gunahkarcasper /corporate-security-and-protection-protocolstextn<1K0 likes2 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.