CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01h8st6ptv /turkey-all-universitiesCertainly! Here’s the dataset description in Markdown format: All Universities in Turkey Dataset Description This dataset contains detailed information about various universities. Each record represents a single university and includes attributes such as the university's name, type, city, website, address, logo URL, and a button for accessing additional details. This data is typically extracted from a web page listing universities. Fields 1. id… See the full description on the dataset page: https://huggingface.co/datasets/h8st6ptv/turkey-all-universities.imagen<1K2 likes22k downloads2y agoHugging Face02turing-motors /Cauldron-JA Dataset Card for The Cauldron-JA Dataset description The Cauldron-JA is a Vision Language Model dataset that translates 'The Cauldron' into Japanese using the DeepL API. The Cauldron is a massive collection of 50 vision-language datasets (training sets only) that were used for the fine-tuning of the vision-language model Idefics2. To create a Japanese Vision Language Dataset, datasets related to OCR, coding, and graphs were excluded because translating them into Japanese… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Cauldron-JA.textvisual-question-answering1M<n<10M9 likes18k downloads2y agoHugging Face03tascib /turkish-llm-dataset Turkish Pretraining Corpus Dataset Description This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication. It is intended for the development, training, and evaluation of Turkish language models. This dataset was prepared as part of a capstone project conducted by a group of students from Sabancı… See the full description on the dataset page: https://huggingface.co/datasets/tascib/turkish-llm-dataset.text100M<n<1B15 likes11k downloads5mo agoHugging Face04tom-gibbs /multi-turn_jailbreak_attack_datasets Multi-Turn Jailbreak Attack Datasets Description This dataset was created to compare single-turn and multi-turn jailbreak attacks on large language models (LLMs). The primary goal is to take a single harmful prompt and distribute the harm over multiple turns, making each prompt appear harmless in isolation. This approach is compared against traditional single-turn attacks with the complete prompt to understand their relative impacts and failure modes. The key feature of… See the full description on the dataset page: https://huggingface.co/datasets/tom-gibbs/multi-turn_jailbreak_attack_datasets.1K<n<10K13 likes11k downloads2y agoHugging Face05serda-dev /turkish-raw-text-cleaned Turkish Raw Text Cleaned turkish-raw-text-cleaned, Türkçe dil modeli çalışmaları için hazırlanmış temizlenmiş ham metin veri kümesidir. Veri kümesi, turkish-nlp-suite çatısı altında yayımlanan Türkçe metin kaynaklarının temizlenmesi, filtrelenmesi ve model eğitimine daha uygun hale getirilmesiyle oluşturulmuştur. Bu çalışma özellikle Türkçe LLM ön-eğitimi, continual pre-training (CPT), tokenizer analizi, embedding modeli eğitimi, alan bağımsız Türkçe metin modelleme ve veri… See the full description on the dataset page: https://huggingface.co/datasets/serda-dev/turkish-raw-text-cleaned.text-generation1M<n<10M0 likes11k downloads3mo agoHugging Face06GEM /wiki_auto_asset_turk Dataset Card for GEM/wiki_auto_asset_turk Link to Main Data Card You can find the main data card on the GEM Website. Dataset Summary WikiAuto is an English simplification dataset that we paired with ASSET and TURK, two very high-quality evaluation datasets, as test sets. The input is an English sentence taken from Wikipedia and the target a simplified sentence. ASSET and TURK contain the same test examples but have references that are simplified in different… See the full description on the dataset page: https://huggingface.co/datasets/GEM/wiki_auto_asset_turk.text100K<n<1M8 likes9.4k downloads2y agoHugging Face07otoearth /otoSpeech-full-duplex-turn-104hgated Dataset Card for otoSpeech-full-duplex-turn-104h Contact Website: https://oto.earthEmail: agent@oto.earth Dataset Summary otoSpeech-full-duplex-turn-104h is an English, full-duplex conversational speech dataset for research on turn-taking and related spoken-dialogue phenomena. It contains 420 two-speaker conversations totaling approximately 104.94 hours. Each conversation includes time-aligned, channel-separated audio, a stereo combined recording… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-turn-104h.audioaudio-to-audio1K<n<10K12 likes7.1k downloads27d agoHugging Face08t2ance /atlas-32-turn-level-actor-critic 32. A turn-level actor-critic derived from the value of computation 1. Question and links Read this first. The reading copy of this directory is t2ance/atlas-experiments under 32-turn-level-actor-critic/; the saved training steps and the per-token training arrays are only in the Hugging Face repository t2ance/atlas-32-turn-level-actor-critic. Does a critic that predicts the return at the start of each turn, and is supervised there alone, learn on the… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-32-turn-level-actor-critic.0 likes5.2k downloads47m agoHugging Face09alan-turing-institute /turing-synthetic-radar-datasetgated The Turing Synthetic Radar Dataset (TSRD) Dataset Summary The Turing Synthetic Radar Dataset is the first publicly available, comprehensively simulated pulse train dataset designed for radar pulse deinterleaving research. It provides a large-scale benchmark for developing and evaluating electronic warfare (EW) and signal intelligence (SIGINT) applications, enabling researchers to address the critical challenge of separating interleaved radar pulses from multiple… See the full description on the dataset page: https://huggingface.co/datasets/alan-turing-institute/turing-synthetic-radar-dataset.1B<n<10B43 likes4.9k downloads4mo agoHugging Face10turandai /gaussian-surfels-dtuimage1K<n<10K2 likes4.4k downloads2y agoHugging Face11Hoshipu /behavior-1k-mp-collected-turning-on-radio BEHAVIOR-1K MP-Collected — turning_on_radio Combined dataset for BEHAVIOR-1K task 0 (turning_on_radio): 1154 success demos + 846 failure demos collected by a hybrid motion-planner + X-VLA policy pipeline on instances 301–700 (private test set, 400 instances × 5 episodes) 200 success demos from the original BEHAVIOR-1K teleoperated dataset (behavior-1k/2025-challenge-demos), merged into success/ Total: 1354 success + 846 failure = 2200 episodes (~157 GB). Layout… See the full description on the dataset page: https://huggingface.co/datasets/Hoshipu/behavior-1k-mp-collected-turning-on-radio.videorobotics10K<n<100K0 likes3.8k downloads4mo agoHugging Face12zinderud /risale-sohbet-turkish-2audio1K<n<10K0 likes3.6k downloads1y agoHugging Face13altaidevorg /fineweb-2-turkish-categorized What is this THis is the categorized version of the Turkish subset of the fineweb-2 dataset. It is an ongoing effort, and the details will be added soon with the rest of the dataset. tabular10M<n<100M15 likes3.6k downloads2y agoHugging Face14turkish-nlp-suite /BellaTurca Dataset Card for BellaTurca BellaTurca is the first large-scale Turkish corpus collection for training Turkish language models. The total size is around 245GB and 30 billion words. BellaTurca's focus is high quality, diversity as well as the size. This collection is made up of five datasets: AkademikDerlem, OzenliDerlem, ForumSohbetleri, Temiz OSCAR and Temiz mC4. Originally there was a book corpus included, but it is excluded due to containing copyrighted material. AkademikDerlem… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/BellaTurca.text10M<n<100M17 likes3.4k downloads7mo agoHugging Face15mundo-ai /turn-benchmark-devgated TurnBench - Dev Set TurnBench is a benchmark for evaluating conversational turn-taking: end-of-turn and interruption detection on real annotated two-speaker conversations. This repository contains the development split: 38 English conversations, about 7.3 hours of audio, packaged as one row per conversation. Each row contains two time-aligned per-speaker audio streams plus three independent annotator tracks per speaker.… See the full description on the dataset page: https://huggingface.co/datasets/mundo-ai/turn-benchmark-dev.audiovoice-activity-detectionn<1K8 likes2.7k downloads1mo agoHugging Face16CharlieLLL /SWEbench-Verified-eval150-turn24to12-8ckpts-2repeats-multicoordinator-w64-20260921 Coding eval150: 24→12 history and checkpoint screening Primary goal: highest absolute accuracy. All selected outcomes and physical attempts are retained. Code: https://github.com/ys-2020/miles/commit/65169383a9bbec29fc138c009d872032bc5ea084 Public evidence archive: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-turn24to12-8ckpts-2repeats-multicoordinator-w64-20260921 Priority Workers Orchestrator History Independent full150 runs A 8 candidates… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-turn24to12-8ckpts-2repeats-multicoordinator-w64-20260921.0 likes2.5k downloads2d agoHugging Face17ioi-leaderboard /ioi-eval-openrouter_openai_gpt-3.5-turbotextn<1K0 likes2.2k downloads2y agoHugging Face18ioi-leaderboard /ioi-eval-dummy-openrouter_openai_gpt-3.5-turbotextn<1K0 likes2.2k downloads2y agoHugging Face19ioi-leaderboard /ioi-eval-openrouter_openai_gpt-3.5-turbo-texttextn<1K0 likes2.2k downloads2y agoHugging Face20ioi-leaderboard /ioi-eval-openrouter_openai_gpt-3.5-turbo-new-prompttextn<1K0 likes2.2k downloads2y agoHugging Face21ioi-leaderboard /ioi-eval-openrouter_openai_gpt-3.5-turbo-prompt-mem-limittextn<1K0 likes2.2k downloads2y agoHugging Face22TurkuNLP /register_oscar Dataset Card for register_oscar Dataset Summary The Register Oscar dataset is a multilingual dataset, containing languaegs from the Oscar dataset that have been tagged with register information. 8 main-level registers: Narrative (NA) Informational Description (IN) Opinion (OP) Interactive Discussion (ID) How-to/Instruction (HI) Informational Persuasion (IP) Lyrical (LY) Spoken (SP) For further description of the labels, see (Douglas Biber and Jesse Egbert. 2018.… See the full description on the dataset page: https://huggingface.co/datasets/TurkuNLP/register_oscar.text1M<n<10M5 likes2.1k downloads3y agoHugging Face23Turki-Alshuaibi /haris-weapon-detection-dataset-curatedimage1 likes2k downloads4mo agoHugging Face24RoboCOIN /Cobot_Magic_turn_on_the_desk_lampgated Cobot_Magic_turn_on_the_desk_lamp 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: agilex_cobot_decoupled_magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: pressbutton 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_turn_on_the_desk_lamp.tabularrobotics100K<n<1M0 likes2k downloads9mo agoHugging Face25turing-motors /Japan-Open-Driving-Dataset-Sample Japan Open Driving Dataset Sample Overview This repository contains a sample subset of the Japan Open Driving Dataset, a large-scale autonomous driving dataset comprising over 100 hours of driving data collected in Tokyo, Japan. The data is stored in nuScenes format and can be loaded with the nuscenes-devkit. In addition to sensor data and 3D annotations, this dataset includes virtual captioned data for training Vision-Language-Model (VLM) and Vision-Language-Action (VLA)… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Japan-Open-Driving-Dataset-Sample.image10K<n<100K5 likes1.9k downloads6mo agoHugging Face26pipecat-ai /smart-turn-data-v3.2-trainTraining dataset for Smart Turn v3.2. Thank you to the following contributors whose audio samples are included in this dataset: The Pipecat team Liva AI: https://www.theliva.ai/ Midcentury: https://www.midcentury.xyz/ MundoAI: https://mundoai.world/ Also, thank you to the following people for the CC-0 background noise sample data which has been used in this dataset: https://freesound.org/people/4team/sounds/214995/ https://freesound.org/people/tomhannen/sounds/698090/… See the full description on the dataset page: https://huggingface.co/datasets/pipecat-ai/smart-turn-data-v3.2-train.audio100K<n<1M11 likes1.9k downloads9mo agoHugging Face27ThaiSyntheticQA /WangchanThaiInstruct_Multi-turn_Conversation_Dataset WangchanThaiInstruct Multi-turn Conversation Dataset We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language. Citation Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633 or BibTeX @dataset{thammaleelakul_2024_13132633, author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.texttext-generation1K<n<10K1 likes1.9k downloads2y agoHugging Face28Merserk /Krea-2-Turbo-Checkpoint-Format-Benchmark Krea 2 Turbo ComfyUI Format Fidelity Benchmark This release is a paired, deterministic comparison of eight Krea 2 Turbo checkpoint formats in ComfyUI: BF16, FP8 Scaled, INT8 ConvRot, MXFP8, NVFP4, INT4 ConvRot W4A4, GGUF Q8_0, and GGUF Q4_K_M. It contains 240 scored 1024×1024 images, saved float32 decoded tensors and final latents, every denoising trajectory, raw metric tables, telemetry, statistical comparisons, and reproduction code. Main result BF16 is the… See the full description on the dataset page: https://huggingface.co/datasets/Merserk/Krea-2-Turbo-Checkpoint-Format-Benchmark.imagetext-to-imagen<1K5 likes1.7k downloads2mo agoHugging Face29RoboCOIN /Cobot_Magic_turn_off_the_desk_lampgated Cobot_Magic_turn_off_the_desk_lamp 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: agilex_cobot_decoupled_magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: pressbutton 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_turn_off_the_desk_lamp.tabularrobotics100K<n<1M0 likes1.7k downloads9mo agoHugging Face30moganai /turkishfineweb2-cleaned TurkishFineweb2-Cleaned A Turkish web corpus derived from the Turkish (tur_Latn) subset of FineWeb-2, augmented with an additional quality-classification layer and a near-duplicate removal pass. 📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum Source FineWeb-2 is a large-scale, multilingual web corpus built from Common Crawl. This dataset covers the Turkish (tur_Latn) portion of FineWeb-2, spanning the… See the full description on the dataset page: https://huggingface.co/datasets/moganai/turkishfineweb2-cleaned.tabulartext-generation10M<n<100M5 likes1.7k downloads1d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.