datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
risale-sohbet-turkish-2risale-i-nur-sohbet
Risale-i Nur Sohbet
Prof. Dr. Şener Dilek’ten izin alındı.
Türkçe
Risale-i Nur sohbetlerini ses, ham ASR metni ve zaman hizalı segmentler hâlinde
birlikte sunan bağımsız bir veri kümesidir. İlk sürüm izinli ve doğrulanmış
sohbetleri içerir; kitap metni, grounded, çok dilli veya kitap seslendirme veri
kümelerine karıştırılmaz.
Kapsam
2095 sohbet, 954.66 saat 16 kHz mono FLAC ses
Aynı derslerin ölçülmüş 48 kHz kalite katmanı; 786 derste
seçici… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-i-nur-sohbet.dolmino-dclmdolmino-mix-1124starcoderdatasts-sohu20212021搜狐校园文本匹配算法大赛数据集TwoHopFactThis is the dataset introduced in the paper Do Large Language Models Latently Perform Multi-Hop Reasoning?.
Code: https://github.com/google-deepmind/latent-multi-hop-reasoning.
german-sohee-synthetic-tts-24k
Flevi Restiti Vici
Bis uns der Weg nur noch nach vorn blieb
This repository contains a synthetic German audiobook and text-to-speech training dataset based on the original novel:
Flevi Restiti ViciBis uns der Weg nur noch nach vorn blieb
The novel was written in German by Maurice Hartmann, who is the author and copyright holder.
Work in progress
This dataset is a work in progress. Future revisions may include corrected transcripts, regenerated audio… See the full description on the dataset page: https://huggingface.co/datasets/Muckylixx/german-sohee-synthetic-tts-24k.casual-conversationRecipePairtrain : 64K pairs
test&validate : 8K pairs
8 : 1 : 1
mini
train : 8K pairs
t&v : 800 pairs
risale-sohbet-turkish
YouTube Transkripsiyon Veri Seti
Veri Yapısı
audio/: MP3 dosyaları
transcripts/: Metin transkripsiyonları
srt/: Altyazı dosyaları
metadata/: Video bilgileri
database.json: Tüm videoların indeksi
Güncelleme Tarihi
2025-03-21
BigEarthNet.txt
BigEarthNet.txt: A Large-Scale Multi-Sensor Image-Text Dataset and Benchmark for Earth Observation
BigEarthNet.txt is a large-scale multi-sensor image–text dataset for Earth observation, designed to advance vision–language learning on remote sensing data. It comprises 464,044 co-registered Sentinel-1 (SAR) and… See the full description on the dataset page: https://huggingface.co/datasets/SohamG2696/BigEarthNet.txt.controlnet-satellite-mapssohl-multidish-yolo-dataset
🍽️ SOHL Multi-Dish Indian Food Detection Dataset
Overview
This dataset contains 377 annotated images of Indian food plates with multiple dishes per image. Designed for training YOLO models to detect and classify multiple food items on a single plate.
Dataset Statistics
Images: 377
Annotations: 377
Classes: 16
Format: YOLOv8 (images + txt annotations)
Created: 2025-08-16
Classes
bread_or_Roti_naan - Chapati, naan, roti, paratha, and other Indian… See the full description on the dataset page: https://huggingface.co/datasets/SohlHealth/sohl-multidish-yolo-dataset.Salesforce-xlam-function-calling-60k-splitsuzbek-legal-irslam_stage2_additional_dataUnhelpfulThoughtsSohamGhadge-casual-conversationReformatted version of SohamGhadge/casual-conversation in ShareGPT-like format.
Each row of this dataset represents a complete path from the root message to the last message in the tree.
The first message (root) is from the user and then it alternates with the AI.
Those paths are truncated if they don't end with the AI's turn.
A random system message was added to each row so the AI will chat naturally and casually with the user.
Limitations:
Because the original dataset only contains one… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/SohamGhadge-casual-conversation.llm-eval-benchmark
LLM Evaluation Benchmark
A 1,200-sample curated benchmark dataset for evaluating LLMs on factual accuracy and truthfulness.
Sourced from MMLU and TruthfulQA, cleaned and formatted for the
LLM Evaluation Framework.
Dataset Summary
Split
Samples
Use
train
500
Fine-tuning reference / training baselines
validation
200
Hyperparameter tuning
test
500
Final benchmark — use this for fair comparisons
Total
1,200
Quick… See the full description on the dataset page: https://huggingface.co/datasets/sohaibdevv/llm-eval-benchmark.SOCRATESThis is the dataset introduced in the paper Do Large Language Models Latently Perform Multi-Hop Reasoning without Exploiting Shortcuts?.
Code: https://github.com/google-deepmind/latent-multi-hop-reasoning.
patient_doctor_chatbotsohi-genuine
sohi-genuine
Mirror of the exact bmcore v24 holdout subset. Source: https://data.mendeley.com/datasets/mvdb8by7yv/1.
1,968 genuine images match the documented 123 participants x 16 genuine signatures; exclude volunteer-forged signatures.
Contains 1968 real image files. This mirror repackages the media; it does not grant additional rights.
Attribution: Siddhartha Kinkor Bayan, Suhair Warish Rahman, Darshita Kalita, and Smriti Priya Medhi (2026), SOHI, Mendeley Data V1, DOI… See the full description on the dataset page: https://huggingface.co/datasets/34data/sohi-genuine.CAD-experiment-manifests-seed202-vision-qwen3vl32b-v1enhanced-indian-food-classification
Enhanced Indian Food Classification Dataset
A comprehensive dataset for Indian food classification with 15,404 images across 43 classes.
Dataset Structure
dataset/
├── train/ # Training images
├── validation/ # Validation images
└── test/ # Test images
Usage
from datasets import load_dataset
# Load dataset
dataset = load_dataset("SohlHealth/enhanced-indian-food-classification")
# Access splits
train_data = dataset['train']
val_data… See the full description on the dataset page: https://huggingface.co/datasets/SohlHealth/enhanced-indian-food-classification.sohl-multidish-yolo-dataset
🍽️ SOHL Multi-Dish Indian Food Detection Dataset
Overview
This dataset contains 377 annotated images of Indian food plates with multiple dishes per image. Designed for training YOLO models to detect and classify multiple food items on a single plate.
Dataset Statistics
Images: 377
Annotations: 377
Classes: 16
Format: YOLOv8 (images + txt annotations)
Created: 2025-08-16
Classes
bread_or_Roti_naan - Chapati, naan, roti, paratha, and other Indian… See the full description on the dataset page: https://huggingface.co/datasets/devmka/sohl-multidish-yolo-dataset.CAD-experiment-manifests-seed202-vision-qwen3vl32b-coderprompts-v1CAD-experiment-manifests-seed42-vision-qwen3vl32b-v1-strict-coder-v1glassdoor_reviewsCAD-experiment-manifests-seed42-vision-qwen3vl4b-coderprompts-v1
