datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Massive-STEPS-Istanbul
Massive-STEPS-Istanbul
Dataset Summary
Massive-STEPSis a large-scale dataset of semantic trajectories intended for understanding POI check-ins. The dataset is derived from the Semantic Trails Dataset and Foursquare Open Source Places, and includes check-in data from 15 cities across 10 countries. The dataset is designed to facilitate research in various domains, including trajectory prediction, POI recommendation, and urban modeling. Massive-STEPS emphasizes the… See the full description on the dataset page: https://huggingface.co/datasets/CRUISEResearchGroup/Massive-STEPS-Istanbul.Panza-emails
The Panza Emails dataset
This dataset contains collections of emails of three authentic users (david, isabel, and marcus), with personal information (names, places, etc.) replaced by other ones for donor privacy.
Except for these changes, the language of the emails is genuine. The intention of this dataset is to allow researchers to study strategies for text personalization.
The data was donated explicitly for this purpose. This dataset is ethically collected and fully licensed for… See the full description on the dataset page: https://huggingface.co/datasets/ISTA-DASLab/Panza-emails.patents-classified-2106-gpt5-miniECB-FED-speeches
ECB and FED Speeches
This data contains speeches from European Central Bank (ECB) and Federal Reserve (FED) executives, from 1996 to 2025.
Mistral OCR
In addition to the text provided by the Bank of International Settlements (BIS), we also added a new textual column derived extracting information from the source PDF files using Mistral's OCR API. Page breaks are identified with the \n\n---[PAGE_BREAK]---\n\n string.
Turkce-istatistik-benchmark
Türkçe İstatistik Benchmark
Bu veri seti, Toivo0/Turkce-istatistik-reasoning ana veri setiyle fine-tuning'de kesinlikle kullanılmamış 100 soru-cevap çiftinden oluşur. Fine-tune edilmiş modellerin gerçek performansını, eğitimde hiç görmediği sorularla ölçmek için hazırlanmıştır.
Benchmark iki aşamada oluşturulmuştur:
37 soru, ana veri setinden (400 soru) eğitim öncesinde stratified (modül bazında orantılı) olarak ayrılmış bölümdür.
63 soru, benchmark'ı istatistiksel olarak daha… See the full description on the dataset page: https://huggingface.co/datasets/Toivo0/Turkce-istatistik-benchmark.Turkce-istatistik-reasoning
Türkçe İstatistik Muhakeme (Chain-of-Thought) Veri Seti
Türkçe'de istatistik konularında düşünce zinciri (chain-of-thought) içeren, soru-cevap formatında bir fine-tuning veri seti. Her örnek, bir kullanıcı sorusu ve modelin hem iç muhakeme sürecini (thinking) hem de nihai cevabını içeren bir asistan yanıtından oluşan bir conversations listesidir.
Veri Seti Özeti
Toplam örnek sayısı
400
Dil
Türkçe
Kapsanan modül sayısı
7
Soru tipleri
Kavramsal… See the full description on the dataset page: https://huggingface.co/datasets/Toivo0/Turkce-istatistik-reasoning.sentipolc_datasetThis is a Hugging Face Dataset wrapper of the Sentipolc Twitter dataset (Basile et al., 2014)
turkish-legal-terms-dictionary
Turkish Legal Terms Dictionary / Türkçe Hukuki Terimler Sözlüğü
🇹🇷 Türkçe
Bu proje, Türkçe hukuki terimlerin ve anlamlarının bulunduğu kapsamlı bir sözlük içermektedir.
📋 İçerik
Toplam terim sayısı: 3.000+ hukuki terim
Format: CSV
Dil: Türkçe
Kapsam: Genel hukuk terimleri, medeni hukuk, ticaret hukuku, ceza hukuku ve diğer hukuk dalları
📁 Dosya Yapısı
├── turkish-legal-terms-dictionary.csv # Ana sözlük dosyası
└── README.md # Bu… See the full description on the dataset page: https://huggingface.co/datasets/istaken/turkish-legal-terms-dictionary.booking-reviews-it-llm
Dataset Card
An adaptation of "Basile, Pierpaolo, et al. "Overview of the EVALITA 2018 Aspect-based Sentiment Analysis task (ABSITA)." EVALITA Evaluation of NLP and Speech Tools for Italian. CEUR, 2018. 1-10.", meant for fine-tuning generative LLMs.
Already contains training, validation, and test splits.
istat_sezioni_censimentoi-statements
I-Statements
This dataset has axproximently 5,335 I-statements generated by Qwen2.5-7B-Q4_K_M using Ollama.
Stats
Metric
Value
Entries
5,334
Total tokens (GPT2)
36,032
Total words
29,735
Avg. tokens per entry
6.67
Avg. words per entry
5.57
Word range
3–10
Unique vocab (words)
2,237
Unique verbs
252
We used GPT2's tokenizer to find the token count.
Note: The tokens may vary depending on the tokenizer used.
Use Cases
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Harley-ml/i-statements.ateco-deterministic-queries
ATECO 2025 Deterministic Queries
This dataset contains around 5k queries extracted from the official ATECO 2025 classification alongside their relative codes and divisions.
court-rulings-coi
Italian Court Rulings on Conflict of Interest (COI)
This dataset contains 1,343 public Italian court rulings related to Conflict of Interest (COI) matters, spanning from 2012 to 2025.
home-ASS-istant-sharegptateco-augmented-queries
ATECO 2025 Augmented Queries
This dataset contains around 12k queries generated by GPT-4.1-mini relative to the official Italian economic activity classification (ATECO 2025).
Around 10 queries for each ATECO code have been generated using the following system prompt:
system = """Sei un generatore di descrizioni di attività svolte da aziende e professionisti.
Ecco cosa devi fare:
* Ricevi in input una classificazione economica (titolo + dettagli).
* Generi 10 esempi di brevissime… See the full description on the dataset page: https://huggingface.co/datasets/istat-ai/ateco-augmented-queries.istanbul_gecmis_havadurumucovid-tweets-100k
Covid Tweets 100k 🇮🇹
This dataset contains 100,000 Italian Tweets related to the Covid-19 pandemic along with their publication date. All texts are lowercase.
istanbul-isitici-kiralama-ile-her-mevsimde-etkinlik-konforuİstanbul’da yılın her dönemi organizasyon yapmak isteyenler için Kirala360, ısıtıcı kiralama hizmetiyle dış mekan etkinliklerinde sıcak ve konforlu bir atmosfer sağlıyor. Düğün, nişan, doğum günü gibi özel günlerinizi ister açık alanda ister yarı kapalı mekanlarda planlayın, farklı ısıtıcı seçenekleri sayesinde mevsim koşullarını düşünmeden davetlerinizi gerçekleştirebilirsiniz. Palmiye mantar ısıtıcıdan ufo modeline, piramit tasarımlı modern ısıtıcılardan fanlı sistemlere kadar zengin ürün… See the full description on the dataset page: https://huggingface.co/datasets/sociallifefour/istanbul-isitici-kiralama-ile-her-mevsimde-etkinlik-konforu.istanbulguidecb-speechesThis dataset contains speeches from central bankers from 1996 to 2024 (September). It includes data from the Bank for International Settlements (BIP) with a few integrations, including a more complete "Institution" field.
istanbul-qa-datasetIf you’d like to use this dataset, please contact omer.erdaggs@gmail.com and provide your reasons.
