datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkey-all-universitiesCertainly! Here’s the dataset description in Markdown format:
All Universities in Turkey Dataset
Description
This dataset contains detailed information about various universities. Each record represents a single university and includes attributes such as the university's name, type, city, website, address, logo URL, and a button for accessing additional details. This data is typically extracted from a web page listing universities.
Fields
1. id… See the full description on the dataset page: https://huggingface.co/datasets/h8st6ptv/turkey-all-universities.Cauldron-JA
Dataset Card for The Cauldron-JA
Dataset description
The Cauldron-JA is a Vision Language Model dataset that translates 'The Cauldron' into Japanese using the DeepL API. The Cauldron is a massive collection of 50 vision-language datasets (training sets only) that were used for the fine-tuning of the vision-language model Idefics2.
To create a Japanese Vision Language Dataset, datasets related to OCR, coding, and graphs were excluded because translating them into Japanese… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Cauldron-JA.turkish-llm-dataset
Turkish Pretraining Corpus
Dataset Description
This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication. It is intended for the development, training, and evaluation of Turkish language models.
This dataset was prepared as part of a capstone project conducted by a group of students from Sabancı… See the full description on the dataset page: https://huggingface.co/datasets/tascib/turkish-llm-dataset.multi-turn_jailbreak_attack_datasets
Multi-Turn Jailbreak Attack Datasets
Description
This dataset was created to compare single-turn and multi-turn jailbreak attacks on large language models (LLMs). The primary goal is to take a single harmful prompt and distribute the harm over multiple turns, making each prompt appear harmless in isolation. This approach is compared against traditional single-turn attacks with the complete prompt to understand their relative impacts and failure modes. The key feature of… See the full description on the dataset page: https://huggingface.co/datasets/tom-gibbs/multi-turn_jailbreak_attack_datasets.turkish-raw-text-cleaned
Turkish Raw Text Cleaned
turkish-raw-text-cleaned, Türkçe dil modeli çalışmaları için hazırlanmış temizlenmiş ham metin veri kümesidir. Veri kümesi, turkish-nlp-suite çatısı altında yayımlanan Türkçe metin kaynaklarının temizlenmesi, filtrelenmesi ve model eğitimine daha uygun hale getirilmesiyle oluşturulmuştur.
Bu çalışma özellikle Türkçe LLM ön-eğitimi, continual pre-training (CPT), tokenizer analizi, embedding modeli eğitimi, alan bağımsız Türkçe metin modelleme ve veri… See the full description on the dataset page: https://huggingface.co/datasets/serda-dev/turkish-raw-text-cleaned.wiki_auto_asset_turk
Dataset Card for GEM/wiki_auto_asset_turk
Link to Main Data Card
You can find the main data card on the GEM Website.
Dataset Summary
WikiAuto is an English simplification dataset that we paired with ASSET and TURK, two very high-quality evaluation datasets, as test sets. The input is an English sentence taken from Wikipedia and the target a simplified sentence. ASSET and TURK contain the same test examples but have references that are simplified in different… See the full description on the dataset page: https://huggingface.co/datasets/GEM/wiki_auto_asset_turk.otoSpeech-full-duplex-turn-104h
Dataset Card for otoSpeech-full-duplex-turn-104h
Contact
Website: https://oto.earthEmail: agent@oto.earth
Dataset Summary
otoSpeech-full-duplex-turn-104h is an English, full-duplex conversational speech dataset for research on turn-taking and related spoken-dialogue phenomena. It contains 420 two-speaker conversations totaling approximately 104.94 hours. Each conversation includes time-aligned, channel-separated audio, a stereo combined recording… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-turn-104h.atlas-32-turn-level-actor-critic
32. A turn-level actor-critic derived from the value of computation
1. Question and links
Read this first. The reading copy of this directory is t2ance/atlas-experiments under 32-turn-level-actor-critic/; the saved training steps and the per-token training arrays are only in the Hugging Face repository t2ance/atlas-32-turn-level-actor-critic.
Does a critic that predicts the return at the start of each turn, and is supervised there alone, learn on the… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-32-turn-level-actor-critic.turing-synthetic-radar-dataset
The Turing Synthetic Radar Dataset (TSRD)
Dataset Summary
The Turing Synthetic Radar Dataset is the first publicly available, comprehensively simulated pulse train dataset designed for radar pulse deinterleaving research. It provides a large-scale benchmark for developing and evaluating electronic warfare (EW) and signal intelligence (SIGINT) applications, enabling researchers to address the critical challenge of separating interleaved radar pulses from multiple… See the full description on the dataset page: https://huggingface.co/datasets/alan-turing-institute/turing-synthetic-radar-dataset.gaussian-surfels-dtubehavior-1k-mp-collected-turning-on-radio
BEHAVIOR-1K MP-Collected — turning_on_radio
Combined dataset for BEHAVIOR-1K task 0 (turning_on_radio):
1154 success demos + 846 failure demos collected by a hybrid motion-planner + X-VLA policy pipeline on instances 301–700 (private test set, 400 instances × 5 episodes)
200 success demos from the original BEHAVIOR-1K teleoperated dataset
(behavior-1k/2025-challenge-demos), merged into success/
Total: 1354 success + 846 failure = 2200 episodes (~157 GB).
Layout… See the full description on the dataset page: https://huggingface.co/datasets/Hoshipu/behavior-1k-mp-collected-turning-on-radio.risale-sohbet-turkish-2fineweb-2-turkish-categorized
What is this
THis is the categorized version of the Turkish subset of the fineweb-2 dataset.
It is an ongoing effort, and the details will be added soon with the rest of the dataset.
BellaTurca
Dataset Card for BellaTurca
BellaTurca is the first large-scale Turkish corpus collection for training Turkish language models. The total size is around 245GB and 30 billion words. BellaTurca's focus is high quality, diversity as well as the size.
This collection is made up of five datasets: AkademikDerlem, OzenliDerlem, ForumSohbetleri, Temiz OSCAR and Temiz mC4. Originally there was a book corpus included, but it is excluded due to containing copyrighted material.
AkademikDerlem… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/BellaTurca.turn-benchmark-dev
TurnBench - Dev Set
TurnBench is a benchmark for evaluating
conversational turn-taking: end-of-turn and interruption detection on real
annotated two-speaker conversations.
This repository contains the development split: 38 English conversations,
about 7.3 hours of audio, packaged as one row per conversation. Each row contains
two time-aligned per-speaker audio streams plus three independent annotator
tracks per speaker.… See the full description on the dataset page: https://huggingface.co/datasets/mundo-ai/turn-benchmark-dev.SWEbench-Verified-eval150-turn24to12-8ckpts-2repeats-multicoordinator-w64-20260921
Coding eval150: 24→12 history and checkpoint screening
Primary goal: highest absolute accuracy. All selected outcomes and physical attempts are retained.
Code: https://github.com/ys-2020/miles/commit/65169383a9bbec29fc138c009d872032bc5ea084
Public evidence archive: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-turn24to12-8ckpts-2repeats-multicoordinator-w64-20260921
Priority
Workers
Orchestrator
History
Independent full150 runs
A
8 candidates… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-turn24to12-8ckpts-2repeats-multicoordinator-w64-20260921.ioi-eval-openrouter_openai_gpt-3.5-turboioi-eval-dummy-openrouter_openai_gpt-3.5-turboioi-eval-openrouter_openai_gpt-3.5-turbo-textioi-eval-openrouter_openai_gpt-3.5-turbo-new-promptioi-eval-openrouter_openai_gpt-3.5-turbo-prompt-mem-limitregister_oscar
Dataset Card for register_oscar
Dataset Summary
The Register Oscar dataset is a multilingual dataset, containing languaegs from the Oscar dataset that have been tagged with register information.
8 main-level registers:
Narrative (NA)
Informational Description (IN)
Opinion (OP)
Interactive Discussion (ID)
How-to/Instruction (HI)
Informational Persuasion (IP)
Lyrical (LY)
Spoken (SP)
For further description of the labels, see (Douglas Biber and Jesse Egbert. 2018.… See the full description on the dataset page: https://huggingface.co/datasets/TurkuNLP/register_oscar.haris-weapon-detection-dataset-curatedCobot_Magic_turn_on_the_desk_lamp
Cobot_Magic_turn_on_the_desk_lamp
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
pressbutton
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_turn_on_the_desk_lamp.Japan-Open-Driving-Dataset-Sample
Japan Open Driving Dataset Sample
Overview
This repository contains a sample subset of the Japan Open Driving Dataset, a large-scale autonomous driving dataset comprising over 100 hours of driving data collected in Tokyo, Japan.
The data is stored in nuScenes format and can be loaded with the nuscenes-devkit.
In addition to sensor data and 3D annotations, this dataset includes virtual captioned data for training Vision-Language-Model (VLM) and Vision-Language-Action (VLA)… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Japan-Open-Driving-Dataset-Sample.smart-turn-data-v3.2-trainTraining dataset for Smart Turn v3.2.
Thank you to the following contributors whose audio samples are included in this dataset:
The Pipecat team
Liva AI: https://www.theliva.ai/
Midcentury: https://www.midcentury.xyz/
MundoAI: https://mundoai.world/
Also, thank you to the following people for the CC-0 background noise sample data which has been used in this dataset:
https://freesound.org/people/4team/sounds/214995/
https://freesound.org/people/tomhannen/sounds/698090/… See the full description on the dataset page: https://huggingface.co/datasets/pipecat-ai/smart-turn-data-v3.2-train.WangchanThaiInstruct_Multi-turn_Conversation_Dataset
WangchanThaiInstruct Multi-turn Conversation Dataset
We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language.
Citation
Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633
or BibTeX
@dataset{thammaleelakul_2024_13132633,
author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.Krea-2-Turbo-Checkpoint-Format-Benchmark
Krea 2 Turbo ComfyUI Format Fidelity Benchmark
This release is a paired, deterministic comparison of eight Krea 2 Turbo checkpoint formats in ComfyUI: BF16, FP8 Scaled, INT8 ConvRot, MXFP8, NVFP4, INT4 ConvRot W4A4, GGUF Q8_0, and GGUF Q4_K_M. It contains 240 scored 1024×1024 images, saved float32 decoded tensors and final latents, every denoising trajectory, raw metric tables, telemetry, statistical comparisons, and reproduction code.
Main result
BF16 is the… See the full description on the dataset page: https://huggingface.co/datasets/Merserk/Krea-2-Turbo-Checkpoint-Format-Benchmark.Cobot_Magic_turn_off_the_desk_lamp
Cobot_Magic_turn_off_the_desk_lamp
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
pressbutton
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_turn_off_the_desk_lamp.turkishfineweb2-cleaned
TurkishFineweb2-Cleaned
A Turkish web corpus derived from the Turkish (tur_Latn) subset of
FineWeb-2, augmented with an additional quality-classification layer
and a near-duplicate removal pass.
📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum
Source
FineWeb-2 is a
large-scale, multilingual web corpus built from Common Crawl. This dataset
covers the Turkish (tur_Latn) portion of FineWeb-2, spanning the… See the full description on the dataset page: https://huggingface.co/datasets/moganai/turkishfineweb2-cleaned.
