CoolFace
Datasetpublic

Zaneoiji/inneratlas-sefaria

INNER ATLAS — Sefaria quotation corpus 21,564 quotation records in 2 file(s), 7,670 of them marked default-visible. Each record carries its own rights block, verification status and attribution string; the per-record rights field is authoritative, not this page. Licence Public Domain, CC0, CC BY, CC BY-SA. Licence Records Public Domain 14,559 CC BY 2,843 CC0 2,648 CC BY-SA 1,514 Attribution What must be displayed with this text… See the full description on the dataset page: https://huggingface.co/datasets/Zaneoiji/inneratlas-sefaria.

sourceHugging Facecc0-1.0updated 6d agoView on Hugging Face
0likes136downloads
Dataset Card

INNER ATLAS — Sefaria quotation corpus

21,564 quotation records in 2 file(s), 7,670 of them marked default-visible. Each record carries its own rights block, verification status and attribution string; the per-record rights field is authoritative, not this page.

Licence

Public Domain, CC0, CC BY, CC BY-SA.

LicenceRecords
Public Domain14,559
CC BY2,843
CC02,648
CC BY-SA1,514

Attribution

What must be displayed with this text, as recorded by the ingestion that produced it:

Sefaria Jewish canon (licence-filtered)

  • —Source: Sefaria — <https://www.sefaria.org>
  • —Licence: mixed: CC0, CC BY, CC BY-SA, Public Domain
  • —Requires: Texts from Sefaria (sefaria.org). Each version is credited to its own edition and carries that version's licence; licences are not inherited between an original and its translation.
  • —Notes: Every version judged on its own recorded licence; nothing inherits from the base text. NonCommercial versions are excluded, which removes the William Davidson Talmud. Language detected from the text, since Sefaria's language tag is unreliable.

Files and integrity

FileRecordsBytessha256
sefaria_en.jsonl.gz9,4911,128,2213c9790394ef3ad67821401d9b28e8b2f7c538ecbda0293acb7cb8aeded2e8b70
sefaria_he.jsonl.gz12,0731,829,144ce8802ab3240d6aa6776f82b6b1db74f2e813269b4d6ff43e5bb0af1a5c33c25

The digests above are the authority. They are recorded in `corpus/corpus_files_v1.json` and verified on every download, so a file that arrives with records added or removed is refused rather than used.

Provenance

Produced by INNER ATLAS from public dumps and APIs, not by scraping. The ingesters, the rights model and the checks that enforce it are in that repository; DATA_LICENSE.md §3.6 records the licensing position and what it does and does not establish.

Not established here: whether the underlying quoted words are free of third-party copyright. This dataset clears the compilation licence of its provider and that provider's own rights determination. Nothing here is a finding about any individual quotation.