NHLOCAL/project-ben-yehuda
Project Ben-Yehuda Corpus Dataset Summary This dataset contains Hebrew literary texts sourced from Project Ben-Yehuda and structured for NLP, language modeling, text analysis, historical linguistics, and computational Hebrew research. Each record represents one text/work and includes the full text, a fixed source label, and structured metadata from the Project Ben-Yehuda catalogue. Current Version Dataset version: pby-2026.03Source snapshot: Project… See the full description on the dataset page: https://huggingface.co/datasets/NHLOCAL/project-ben-yehuda.
Update Project Ben-Yehuda dataset
Delete old files before upload: Update Project Ben-Yehuda dataset
Update Project Ben-Yehuda dataset
Delete old files before upload: Update Project Ben-Yehuda dataset
Update Project Ben-Yehuda dataset
Delete old files before upload: Update Project Ben-Yehuda dataset
Update README.md
Delete data/pby_dataset.parquet
Update dataset from GitHub Actions
Update dataset from GitHub Actions
Update dataset from GitHub Actions
Update dataset from GitHub Actions
Update README.md
Update README.md
Update dataset from GitHub Actions
Update README.md
Update dataset from GitHub Actions
Update dataset from GitHub Actions
Update README.md
Update README.md
initial commit
