datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gazzetta-ufficiale
Gazzetta Ufficiale 👩🏻⚖️⚖️🏛️📜🇮🇹
La Gazzetta Ufficiale della Repubblica Italiana, quale fonte ufficiale di conoscenza delle norme in vigore in Italia e strumento di diffusione, informazione e ufficializzazione di testi legislativi, atti pubblici e privati, è edita dall’Istituto Poligrafico e Zecca dello Stato e pubblicata in collaborazione con il Ministero della Giustizia, il quale provvede alla direzione e redazione della stessa. L'Istituto Poligrafico e Zecca dello Stato… See the full description on the dataset page: https://huggingface.co/datasets/mii-llm/gazzetta-ufficiale.gazette-hu
Magyar Közlöny — Hungarian official gazette (public-domain legal text)
Legislation, decrees and official notices from Magyar Közlöny,
Hungary's official gazette, cleaned from the source PDFs to raw prose. One row per paragraph,
with issue metadata (number, date, year, era) and a best-effort section (the current act
heading). Text only.
License
CC0-1.0. Under the Hungarian Copyright Act (Szjt., Act LXXVI/1999) §1(4), legislation
and other official documents… See the full description on the dataset page: https://huggingface.co/datasets/lazos/gazette-hu.turkish-resmi-gazete
Turkish Resmi Gazete (Official Gazette) — Full Corpus
Türkiye Resmî Gazete'sinin 2019-01-02'den 2026-07-17'ye kadar yayımlanan tüm sayılarının tam metnini içeren bir derlemdir. Kanunlar, yönetmelikler, Cumhurbaşkanı kararları, tebliğler, Anayasa Mahkemesi kararları, kurul kararları ve atama kararnameleri dahil olmak üzere Resmî Gazete'de o tarih aralığında yayımlanmış her türden resmî belge yer almaktadır.
İçerik ve Kaynak
Resmî Gazete, günlük yayınlarını… See the full description on the dataset page: https://huggingface.co/datasets/moganai/turkish-resmi-gazete.ottoman-place-names-gazetteer
Ottoman Turkish Place Names Gazetteer (Transliteration Dataset)
Dataset Summary
This dataset serves as a specialized parallel corpus for Ottoman Turkish to Modern Turkish Latin script transliteration, focusing specifically on historical place names (toponyms). It is designed to enhance the performance of Large Language Models (LLMs) and OCR post-processing tools in recognizing and correctly transcribing historical geographical entities.
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/ottoman-place-names-gazetteer.gazet-dataset
Gazet Dataset
Synthetic training data for finetuning small language models on geospatial tasks over Overture Maps and Natural Earth parquet datasets.
Tasks
SQL generation (sql/)
Input: user query + fuzzy-matched candidate entities (CSV)Output: DuckDB spatial SQL query
Place extraction (places/)
Input: natural language queryOutput: structured JSON with place names, country codes, and subtypes
Format
Each JSONL row is a conversation in… See the full description on the dataset page: https://huggingface.co/datasets/developmentseed/gazet-dataset.finetranslations-gaz-latn
FineTranslations gaz_Latn
A bounded Oromo (gaz_Latn) to English translation workload derived from
HuggingFaceFW/finetranslations,
packaged for batched offline LLM inference experiments.
The slice keeps rows whose full chat-template input length is
input_tokens <= 10000. No source text is truncated. This preserves a meaningful
single-language FineTranslations subset while removing only the longest tail that would dominate
runtime and exceed the intended benchmark shape.… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-systalyze/finetranslations-gaz-latn.gazal-dataset
Gazal Dataset
This dataset contains Urdu gazals (poetry) with their English translations and generated prompts for training language models.
⚠️ IMPORTANT: This dataset is strictly for educational purposes only.
Dataset Structure
The dataset contains the following columns:
author: Value(dtype='string', id=None)
gazal_text: Value(dtype='string', id=None)
language: Value(dtype='string', id=None)
prompt: Value(dtype='string', id=None)
Usage
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/pathikg/gazal-dataset.
