datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dead-sea-scrolls
Dead Sea Scrolls (DSS) Morphology Dataset
Word-level morphological annotations for the Dead Sea Scrolls collection —
biblical and non-biblical scrolls in Ancient Hebrew, Aramaic, and other
ancient languages — derived from the ETCBC/DSS Text-Fabric corpus. Part of
the NuBerea corpus estate.
License
CC BY-NC 4.0, following the data license metadata distributed with the
upstream Text-Fabric corpus. Non-commercial use only.
Attribution
Source… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/dead-sea-scrolls.zero_scrollsPenguinScrolls
PenguinScrolls: A User-Aligned Fine-Grained Benchmark for Long-Context Language Model Evaluation
Introduction
PenguinScrolls (企鹅卷轴) is a comprehensive benchmark designed to evaluate and enhance the long-text processing capabilities of large language models (LLMs).
Current benchmarks for evaluating long-context language models often rely on synthetic tasks that fail to adequately reflect real user needs, leading to a weak correlation between benchmark scores and actual… See the full description on the dataset page: https://huggingface.co/datasets/Penguin-Scrolls/PenguinScrolls.full-scrolls
full-scrolls
Fiber orientation direction fields for the Vesuvius Challenge Herculaneum scrolls,
organized to match the VC3D volpkg folder structure exactly.
What this is
Each tile contains predicted fiber orientation vectors (normal, vertical, horizontal
components) for 256x256x256 voxel cubes of each scroll's CT volume. These are
intended as inputs to VC3D's Surface Growth fiber-direction optimization slot,
to help the segmentation pipeline trace individual… See the full description on the dataset page: https://huggingface.co/datasets/abundantjoe/full-scrolls.
