heb
Datasets
All datasets matching “heb”HebDB
HebDB
Paper: http://arxiv.org/abs/2407.07566
If you use our datasets, please use the following:
@article{turetzky2024hebdb,
title={HebDB: a Weakly Supervised Dataset for Hebrew Speech Processing},
author={Turetzky, Arnon and Tal, Or and Segal-Feldman, Yael and Dissen, Yehoshua and Zeldes, Ella and Roth, Amit and Cohen, Eyal and Shrem, Yosi and Chernyak, Bronya R and Seleznova, Olga and others},
journal={arXiv preprint arXiv:2407.07566},
year={2024}
}… See the full description on the dataset page: https://huggingface.co/datasets/SLPRL-HUJI/HebDB.hebrew-hrm-corpus
Hebrew HRM-Text Corpus
Training corpus for a Hebrew Hierarchical Reasoning Model, replicating the
sapientinc/HRM-Text-1B recipe
(train-from-scratch, PrefixLM over {condition, instruction, response}, loss on response only).
Schema
Each line: {"condition": "<tags>", "instruction": "...", "response": "..."}.
Condition tags map to special tokens: direct→<|object_ref_start|>, cot→<|object_ref_end|>,
noisy→<|quad_start|>, synth→<|quad_end|> (composite tags… See the full description on the dataset page: https://huggingface.co/datasets/guychuk/hebrew-hrm-corpus.chat-resultsheb-fonts-2font-hh104-hh1heb-connected-composed-4stylehebocrbench-v1
HebOCRBench v1 — Modern Hebrew participant pack
This is the public, fixed evaluation-input pack for the five-track
HebOCRBench 1.0 Modern Hebrew printed-document headline. It contains 34,267
images and opaque input records. It intentionally contains no transcription
gold, original source identifiers, source URLs, private filenames, organizer
ID map, or HMAC key.
The unchanged certified references are withheld from this repository. An
organizer remaps submitted opaque identifiers… See the full description on the dataset page: https://huggingface.co/datasets/ssdataanalysis/hebocrbench-v1.
