datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hebrew-space-restoration-corpus
Restoring Missing Spaces in Scraped Hebrew Social Media
This dataset holds the test corpus used in the 2025 W-Nut paper: Avi Shmidman and Shaltiel Shmidman, "Restoring Missing Spaces in Scraped Hebrew Social Media", The 10th Workshop on Noisy and User-generated Text (W-NUT), 2025.
The corpus consists of ~6,000 Hebrew sentences, sampled from the Hebrew portion of FineWeb-2.
Each row of the dataset contains two fields:
input: The Hebrew sentence with 1-4 spaces randomly removed (see… See the full description on the dataset page: https://huggingface.co/datasets/dicta-il/hebrew-space-restoration-corpus.incomplete_utterance_restoration
Датасет для задачи раскрытия неполных реплик в контексте диалога
Подробное описание задачи "Incomplete Utterance Restoration" можно найти в карточке генеративной модели inkoziev/rugpt_interpreter, которая обучена на аугментированном варианте этого датасета.
В датасете содержатся фрагменты диалогов длиной от 1 до 3 последовательных реплик. Для последней реплики дается ее полный вариант с раскрытыми анафорами, эллипсисами и т.д.
Например, следующий сэмпл:
{
"context":… See the full description on the dataset page: https://huggingface.co/datasets/inkoziev/incomplete_utterance_restoration.
