CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zicsx /mC4-Hindi-Cleaned-3.0 Dataset Card for "mC4-Hindi-Cleaned-3.0" More Information needed text1M<n<10M2 likes4.4k downloads3y agoHugging Face02izumi-lab /mc4-ja Dataset Card for "mc4-ja" More Information needed text10M<n<100M6 likes4k downloads3y agoHugging Face03Finnish-NLP /mc4_3.1.0_fi_cleaned Dataset Card for "mc4_3.1.0_fi_cleaned" More Information needed tabular10M<n<100M0 likes1.5k downloads3y agoHugging Face04izumi-lab /mc4-ja-filter-ja-normal Dataset Card for "mc4-ja-filter-ja-normal" More Information needed text10M<n<100M5 likes915 downloads3y agoHugging Face05zicsx /mC4-Hindi-Cleaned Dataset Card for "mC4-Hindi-Cleaned" More Information needed text1M<n<10M0 likes834 downloads3y agoHugging Face06eduagarcia /mc4-pt MC4-PT MC4-PT is the is the portuguese subset from MC4. MC4 is a multilingual colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the raw version. Deduplicated version is available here. text100M<n<1B2 likes682 downloads3y agoHugging Face07jorgeortizfuentes /mc4_es_cl Dataset Card for "mc4_es_cl" More Information needed text1M<n<10M1 likes569 downloads4y agoHugging Face08bowphs /mc4-fractiontext1M<n<10M0 likes431 downloads2y agoHugging Face09markus583 /mC4-TESTtext100M<n<1B0 likes417 downloads3y agoHugging Face105CD-AI /Viet-Font-mc4-Textimage10K<n<100K3 likes390 downloads2y agoHugging Face11tiennv /english-mc4 Dataset Card for "english-mc4" More Information needed text10M<n<100M0 likes350 downloads3y agoHugging Face12zicsx /mC4-hindi Dataset Card for "mC4-hindi" This dataset is a subset of the mC4 dataset, which is a multilingual colossal, cleaned version of Common Crawl's web crawl corpus. It contains natural text in 101 languages, including Hindi. This dataset is specifically focused on Hindi text, and contains a variety of different types of text, including news articles, blog posts, and social media posts. This dataset is intended to be used for training and evaluating natural language processing models for… See the full description on the dataset page: https://huggingface.co/datasets/zicsx/mC4-hindi.texttext-generation10M<n<100M0 likes342 downloads3y agoHugging Face13zicsx /mC4-Hindi-Cleaned-2.0 Dataset Card for "test" More Information needed text1M<n<10M1 likes120 downloads3y agoHugging Face14tmfi /mc4-jatext10M<n<100M0 likes112 downloads3y agoHugging Face15duckaiml /mc4_310mc4 but in HPC friendly parquet format (32GiB shards) Attribution,license, copyright info: Google and AI^2 for producing and uploading them. text10M<n<100M0 likes80 downloads3y agoHugging Face16Polygl0t /mc4-50k mC4 (50K samples per language) Dataset Summary This is the multilingual version of the C4 dataset. We only keep 50,000 samples per language for all the 108 languages available. This is intended for educational and experimentation purposes. For the full dataset, please refer to the multilingual config of the allenai/c4 dataset. How do I download this? Using 🤗 Datasets from datasets import load_dataset # Portuguese only pt =… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/mc4-50k.text1M<n<10M0 likes72 downloads9mo agoHugging Face17DomainLLM /legal-mc4-mcqstextn<1K0 likes68 downloads1y agoHugging Face18jcmc /ga_mc4_processedtext100K<n<1M0 likes50 downloads5y agoHugging Face19chenghao /mc4_eu_deduptext1M<n<10M0 likes45 downloads5y agoHugging Face20ayganyavuz /mc4_tr_1gb_sampletext100K<n<1M0 likes43 downloads10mo agoHugging Face21mrm8488 /mc4-es-300k-cosmopedia-en-100ktext100K<n<1M0 likes40 downloads2y agoHugging Face22mrm8488 /mc4-es-300ktext100K<n<1M3 likes39 downloads2y agoHugging Face23chenghao /mc4_sw_deduptext100K<n<1M0 likes37 downloads5y agoHugging Face24sma1-rmarud /plat-kor-mc4 PLAT-KOR-MC4 PLAT (Predicting the Legitimacy of punitive Additional Tax) - 한국어 객관식 (A/B/C/D) 데이터셋 설명 한국 조세법 판례 100개를 4지선다형 문제로 구성한 데이터셋입니다. 데이터셋 구조 필드 타입 설명 case_no string 사건 번호 (예: "2011구합2638") case_info string 사건 개요 (당사자, 배경) facts string 사건 사실관계 claims string 원고와 피고의 주장 reasoning string 법원의 법적 판단 decision string 법원의 최종 판결 choices list[string] 4개의 선택지 gt string 정답 (선택지 중 하나) 사용법 아래의 github를 참조.… See the full description on the dataset page: https://huggingface.co/datasets/sma1-rmarud/plat-kor-mc4.textquestion-answeringn<1K0 likes34 downloads8mo agoHugging Face25nhagar /legal-mc4_urls Dataset Card for legal-mc4_urls This dataset provides the URLs and top-level domains associated with training records in joelniklaus/legal-mc4. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/legal-mc4_urls.text1M<n<10M0 likes29 downloads1y agoHugging Face26anti-ai /vi_mc4_biology_wseggated Dataset Card for "vi_mc4_biology_wseg" More Information needed text1M<n<10M1 likes27 downloads3y agoHugging Face27sma1-rmarud /plat-eng-mc4 PLAT-ENG-MC4 PLAT (Predicting the Legitimacy of punitive Additional Tax) - English 4-Choice Multiple Choice Dataset Description This dataset contains 100 Korean tax law cases translated to English, formatted as 4-choice multiple choice questions. Dataset Structure Field Type Description case_no string Case identifier (e.g., "2011guhap2638") case_info string Summary of the case (parties, background) facts string Detailed facts of the case claims… See the full description on the dataset page: https://huggingface.co/datasets/sma1-rmarud/plat-eng-mc4.textquestion-answeringn<1K0 likes22 downloads8mo agoHugging Face28ttaront /iz_mc4_jptext10K<n<100K0 likes21 downloads2y agoHugging Face29riimaru /MC4_2jttext1M<n<10M0 likes20 downloads3mo agoHugging Face30oriki101 /mc4_ja_text_volume_annotatted_dataこのデータセットはmC4の日本語データに対し、文の割合に応じて1~5段階で人手で評価したものです。文の割合はデータセットのscoreに格納しています。 文の割合が20%以下 文の割合が20~40% 文の割合が40~60% 文の割合が60~80% 文の割合が80~100% データ量は500件と少ないですが、mC4のゴミデータ削除に役立てればと思います。 tabularn<1K0 likes15 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.