hebrew
Datasets
All datasets matching “hebrew”hebrew-hrm-corpus
Hebrew HRM-Text Corpus
Training corpus for a Hebrew Hierarchical Reasoning Model, replicating the
sapientinc/HRM-Text-1B recipe
(train-from-scratch, PrefixLM over {condition, instruction, response}, loss on response only).
Schema
Each line: {"condition": "<tags>", "instruction": "...", "response": "..."}.
Condition tags map to special tokens: direct→<|object_ref_start|>, cot→<|object_ref_end|>,
noisy→<|quad_start|>, synth→<|quad_end|> (composite tags… See the full description on the dataset page: https://huggingface.co/datasets/guychuk/hebrew-hrm-corpus.Hebrew-Mila collection of pdfs about militery and defence hebrew only
chat-resultshebrew_pretrain_v1_4baseData_processedmacula-hebrew-syntax
NuBerea MACULA Hebrew Syntax Trees (OT)
Full syntactic tree annotation of the Hebrew Bible from the MACULA Hebrew Linguistic Dataset, packaged as relational tables for computational biblical studies. The dataset covers word-level linguistic annotation (morphology, glosses, lexical semantics), sentence segmentation, and hierarchical syntactic structure (clauses and phrases with their roles and containment relations) over the Westminster Leningrad Codex base text.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/macula-hebrew-syntax.hebrew_speech_coursera
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
{'audio': {'path':… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/hebrew_speech_coursera.
