CoolFace
Datasetpublic

Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_pdf

llm-jp-corpus-v4 — ja_sip_comprehensive_pdf Mirror of the ja/ja_sip_comprehensive_pdf sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_sip_comprehensive_pdf Files: 156 × jsonl.gz (39.1 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_pdf.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes496downloads
Dataset Card

llm-jp-corpus-v4 — ja_sip_comprehensive_pdf

Mirror of the ja/ja_sip_comprehensive_pdf sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII).

  • Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
  • Sub-corpus: ja_sip_comprehensive_pdf
  • Files: 156 × jsonl.gz (39.1 GB compressed)
  • Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields).

Directory layout mirrors the upstream repository.

License

CC BY 4.0 — inherited from the upstream sub-corpus. Attribution goes to the LLM-jp Corpus Building WG and to the original data providers listed in the upstream README, which remains the authoritative statement of terms for this data.