datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gazette-hu
Magyar Közlöny — Hungarian official gazette (public-domain legal text)
Legislation, decrees and official notices from Magyar Közlöny,
Hungary's official gazette, cleaned from the source PDFs to raw prose. One row per paragraph,
with issue metadata (number, date, year, era) and a best-effort section (the current act
heading). Text only.
License
CC0-1.0. Under the Hungarian Copyright Act (Szjt., Act LXXVI/1999) §1(4), legislation
and other official documents… See the full description on the dataset page: https://huggingface.co/datasets/lazos/gazette-hu.turkish-resmi-gazete
Turkish Resmi Gazete (Official Gazette) — Full Corpus
Türkiye Resmî Gazete'sinin 2019-01-02'den 2026-07-17'ye kadar yayımlanan tüm sayılarının tam metnini içeren bir derlemdir. Kanunlar, yönetmelikler, Cumhurbaşkanı kararları, tebliğler, Anayasa Mahkemesi kararları, kurul kararları ve atama kararnameleri dahil olmak üzere Resmî Gazete'de o tarih aralığında yayımlanmış her türden resmî belge yer almaktadır.
İçerik ve Kaynak
Resmî Gazete, günlük yayınlarını hem düz… See the full description on the dataset page: https://huggingface.co/datasets/moganai/turkish-resmi-gazete.ottoman-place-names-gazetteer
Ottoman Turkish Place Names Gazetteer (Transliteration Dataset)
Dataset Summary
This dataset serves as a specialized parallel corpus for Ottoman Turkish to Modern Turkish Latin script transliteration, focusing specifically on historical place names (toponyms). It is designed to enhance the performance of Large Language Models (LLMs) and OCR post-processing tools in recognizing and correctly transcribing historical geographical entities.
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/ottoman-place-names-gazetteer.gazelle_benchmark
Gazelle: An Instruction Dataset for Arabic Writing Assistance
Writing has long been considered a hallmark of human intelligence and remains a pinnacle task for artificial intelligence (AI) due to the intricate cognitive processes involved. Recently, rapid advancements in generative AI, particularly through the development of Large Language Models (LLMs), have significantly transformed the landscape of writing assistance. However, underrepresented languages like Arabic encounter… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/gazelle_benchmark.gazet-dataset
Gazet Dataset
Synthetic training data for finetuning small language models on geospatial tasks over Overture Maps and Natural Earth parquet datasets.
Tasks
SQL generation (sql/)
Input: user query + fuzzy-matched candidate entities (CSV)Output: DuckDB spatial SQL query
Place extraction (places/)
Input: natural language queryOutput: structured JSON with place names, country codes, and subtypes
Format
Each JSONL row is a conversation in… See the full description on the dataset page: https://huggingface.co/datasets/developmentseed/gazet-dataset.GazeXplain
GazeXplain: Learning to Predict Natural Language Explanations of Visual Scanpaths
GazeXplain is a novel study that goes beyond predicting where people look; it demands models to explain them in natural language, weaving a narrative thread that connects fixations to their underlying meaning.
The repository contains the datasets of the explanations of visual scanpaths in three different scanpath datasets (OSIE, AiR-D, COCO-Search18).
Free-viewing: the prediction of scanpath for… See the full description on the dataset page: https://huggingface.co/datasets/chenxy99/GazeXplain.
