datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mC4-Hindi-Cleaned-3.0
Dataset Card for "mC4-Hindi-Cleaned-3.0"
More Information needed
mc4-ja
Dataset Card for "mc4-ja"
More Information needed
mc4_3.1.0_fi_cleaned
Dataset Card for "mc4_3.1.0_fi_cleaned"
More Information needed
mc4-ja-filter-ja-normal
Dataset Card for "mc4-ja-filter-ja-normal"
More Information needed
mC4-Hindi-Cleaned
Dataset Card for "mC4-Hindi-Cleaned"
More Information needed
mc4-pt
MC4-PT
MC4-PT is the is the portuguese subset from MC4.
MC4 is a multilingual colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org".
This is the raw version. Deduplicated version is available here.
mc4_es_cl
Dataset Card for "mc4_es_cl"
More Information needed
mc4-fractionmC4-TESTViet-Font-mc4-Textenglish-mc4
Dataset Card for "english-mc4"
More Information needed
mC4-hindi
Dataset Card for "mC4-hindi"
This dataset is a subset of the mC4 dataset, which is a multilingual colossal, cleaned version of Common Crawl's web crawl corpus. It contains natural text in 101 languages, including Hindi. This dataset is specifically focused on Hindi text, and contains a variety of different types of text, including news articles, blog posts, and social media posts.
This dataset is intended to be used for training and evaluating natural language processing models for… See the full description on the dataset page: https://huggingface.co/datasets/zicsx/mC4-hindi.mC4-Hindi-Cleaned-2.0
Dataset Card for "test"
More Information needed
mc4-jamc4_310mc4 but in HPC friendly parquet format (32GiB shards)
Attribution,license, copyright info: Google and AI^2 for producing and uploading them.
mc4-50k
mC4 (50K samples per language)
Dataset Summary
This is the multilingual version of the C4 dataset. We only keep 50,000 samples per language for all the 108 languages available. This is intended for educational and experimentation purposes. For the full dataset, please refer to the multilingual config of the allenai/c4 dataset.
How do I download this?
Using 🤗 Datasets
from datasets import load_dataset
# Portuguese only
pt =… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/mc4-50k.legal-mc4-mcqsga_mc4_processedmc4_eu_dedupmc4_tr_1gb_samplemc4-es-300k-cosmopedia-en-100kmc4-es-300kmc4_sw_dedupplat-kor-mc4
PLAT-KOR-MC4
PLAT (Predicting the Legitimacy of punitive Additional Tax) - 한국어 객관식 (A/B/C/D)
데이터셋 설명
한국 조세법 판례 100개를 4지선다형 문제로 구성한 데이터셋입니다.
데이터셋 구조
필드
타입
설명
case_no
string
사건 번호 (예: "2011구합2638")
case_info
string
사건 개요 (당사자, 배경)
facts
string
사건 사실관계
claims
string
원고와 피고의 주장
reasoning
string
법원의 법적 판단
decision
string
법원의 최종 판결
choices
list[string]
4개의 선택지
gt
string
정답 (선택지 중 하나)
사용법
아래의 github를 참조.… See the full description on the dataset page: https://huggingface.co/datasets/sma1-rmarud/plat-kor-mc4.legal-mc4_urls
Dataset Card for legal-mc4_urls
This dataset provides the URLs and top-level domains associated with training records in joelniklaus/legal-mc4. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/legal-mc4_urls.vi_mc4_biology_wseg
Dataset Card for "vi_mc4_biology_wseg"
More Information needed
plat-eng-mc4
PLAT-ENG-MC4
PLAT (Predicting the Legitimacy of punitive Additional Tax) - English 4-Choice Multiple Choice
Dataset Description
This dataset contains 100 Korean tax law cases translated to English, formatted as 4-choice multiple choice questions.
Dataset Structure
Field
Type
Description
case_no
string
Case identifier (e.g., "2011guhap2638")
case_info
string
Summary of the case (parties, background)
facts
string
Detailed facts of the case
claims… See the full description on the dataset page: https://huggingface.co/datasets/sma1-rmarud/plat-eng-mc4.iz_mc4_jpMC4_2jtmc4_ja_text_volume_annotatted_dataこのデータセットはmC4の日本語データに対し、文の割合に応じて1~5段階で人手で評価したものです。文の割合はデータセットのscoreに格納しています。
文の割合が20%以下
文の割合が20~40%
文の割合が40~60%
文の割合が60~80%
文の割合が80~100%
データ量は500件と少ないですが、mC4のゴミデータ削除に役立てればと思います。
