CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Eccentricity /bashbench2 BashBench2 A successor in spirit to the original BashBench, this dataset is intended for high-stakes agentic control research using current and near-future frontier models. The code required to set up and run these tasks is located in ControlArena. textn<1K1 likes522 downloads11mo agoHugging Face02AigizK /bashkort_commands_omnivoice Bashkort Commands OmniVoice Partial eleven-label command snapshot generated with k2-fsa/OmniVoice using the same cross-lingual voice-cloning recipe as AigizK/homai_wake_word_omnivoice. Generation was stopped at the user's request after 41,525 complete reference groups had been committed. For every included reference row from the train split of: bond005/sova_rudevices the dataset contains one recording of every command: Айвика — Russian Айвикә — Bashkir Айһылыу — Bashkir… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/bashkort_commands_omnivoice.audioaudio-classification100K<n<1M0 likes353 downloads2mo agoHugging Face03failed09 /bashkir-frequency-index Bashkir Frequency Index v11.5 Word-frequency index for Bashkir, computed over a large monolingual Bashkir-language dataset, for NLP, spellchecking and lexical research. Overview Word-frequency index for the Bashkir language computed over a large monolingual Bashkir-language dataset. Non-Bashkir admixture, borrowed vocabulary and scanning artifacts were reduced with automated language filtering. The public configuration (count ≥ 3) is the recommended default;… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-frequency-index.tabulartext-classification1M<n<10M0 likes240 downloads3d agoHugging Face04emirkaanozdemr /bash_command_data_6K 📦 Bash Command Dataset v1 A high-quality dataset of natural language instructions paired with their equivalent Bash commands, designed for training and fine-tuning large language models (LLMs) that translate English tasks into shell commands. This dataset is ideal for researchers, developers, and machine learning engineers interested in natural language to Bash command translation, command-line automation, and building intelligent terminal assistants. 📁 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/emirkaanozdemr/bash_command_data_6K.texttext-generation1K<n<10K8 likes206 downloads6mo agoHugging Face05failed09 /bashkir-multilingual-phrasebooks Bashkir-Russian Phrasebook Corpus Edited Bashkir-Russian words, expressions and conversational phrases from university phrasebooks, annotated by entry type. Overview Edited Bashkir-Russian pairs derived from the original bashkorttele/trilingual-parallel-phrasebooks-bgpu dataset, published by Bashkorttele from phrasebooks of M. Akmulla Bashkir State Pedagogical University. The cleaned configuration is the deduplicated default; reviewed is the edited edition… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-multilingual-phrasebooks.texttranslation10K<n<100K0 likes158 downloads7d agoHugging Face06AigizK /bashkort_voice Bashkort Voice 🇬🇧 English Version Dataset Description This is a synthetic Bashkir audio dataset generated using the OmniVoice model. It is designed to expand the availability of spoken data for the Bashkir language. Data Preparation Process The dataset was constructed through a cross-lingual voice cloning and generation process, using the following methodology: Target Text: Bashkir sentences were extracted from the AigizK/bashkir-russian-parallel-corpora dataset.… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/bashkort_voice.audioautomatic-speech-recognition100K<n<1M2 likes156 downloads5mo agoHugging Face07failed09 /bashkir-wikipedia-parallel Bashkir-Russian Wikipedia Parallel Corpus Sentence-level Bashkir-Russian parallel text from Wikipedia, scored and filtered for machine translation. Overview Sentence-level Bashkir-Russian parallel dataset extracted from the corresponding Bashkir and Russian Wikipedia dumps dated 2026-08-01. Candidate pairs are scored for semantic alignment with multilingual sentence encoders (Meta LASER3, Google LaBSE) and the in-domain Bashkir-Russian Pair Scorer. The filtered… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-parallel.tabulartranslation100K<n<1M0 likes151 downloads7d agoHugging Face08AigizK /bashkort_tts_dataset Bashkort TTS Dataset The largest open dataset for speech synthesis in the Bashkir language — featuring multi-speaker recordings and speaking styles. 📊 Dataset Overview Total audio files: 62,852 Speakers: 7 female, 1 male Speaking styles: friendly, question, neutral Languages: Bashkir Format: MP3 audio + transcription text 🎙 How It Was Collected Initial recording: A female voice actor recorded ~15 hours of speech in Bashkir. Voice cloning: Using ElevenLabs… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/bashkort_tts_dataset.audiotext-to-speech10K<n<100K3 likes147 downloads1y agoHugging Face09Bashifu /uav-fault-symptom-reports UAV Fault Symptom Reports A synthetic dataset of UAV flight telemetry paired with operator-style symptom reports written by a language model. Each row is one five-second window of a flight: 20 telemetry channels, the fault class, a severity derived from simulated consequences, and a one-sentence report. split rows flights model-written reports unique reports benchmark 10,500 2,100 82.0% 79.7% challenge 3,500 700 88.1% 91.0% benchmark is balanced across seven… See the full description on the dataset page: https://huggingface.co/datasets/Bashifu/uav-fault-symptom-reports.imagetext-classification10K<n<100K0 likes117 downloads1d agoHugging Face10metuKKhud /bashqort-raw Bashqort Raw Corpus Description This dataset contains raw Bashkir text collected for continual training of large language models (LLMs). It is part of the project "Adapting Open-Source LLMs for the Bashkir Language", which aims to evaluate adaptation methods proposed by LlamaTurk (Toraman, 2024) and Persian adaptation (Mahdizadeh Sani et al., 2024). The corpus is assembled from multiple sources to provide a diverse linguistic foundation for language modeling.… See the full description on the dataset page: https://huggingface.co/datasets/metuKKhud/bashqort-raw.text100K<n<1M0 likes105 downloads4mo agoHugging Face11failed09 /bashkir-ngram-index Bashkir Word N-gram Index v11.5 Exact within-sentence word n-gram counts for Bashkir: unigrams, bigrams and trigrams for spellchecking, OCR post-processing and lightweight language modelling. Overview Exact word n-gram counts derived from a monolingual Bashkir-language dataset. The release provides unigram, bigram and trigram indexes for corpus processing, spellchecking, OCR post-processing, autocomplete and lightweight language-model experiments. The unigrams… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-ngram-index.tabulartext-classification10M<n<100M0 likes105 downloads3d agoHugging Face12DCAgent /rl__40GPU_base_32b__exp_rpt_nemotron-bash__Qwen3-32Btext10K<n<100K0 likes103 downloads7mo agoHugging Face13DCAgent2 /DCAgent2_terminal_bench_2_DCAgent_bash_textbook_tasks_traces_20251123_000222textn<1K0 likes95 downloads10mo agoHugging Face14GunA-SD /bash_codeThis dataset is a collection of bash programs from various GitHub repositories and open source projects. The dataset might contain harmful code. texttext-generation100K<n<1M9 likes86 downloads2y agoHugging Face15AigizK /bashkir-russian-parallel-corpora Dataset Card for "bashkir-russian-parallel-corpora" How the dataset was assembled. find the text in two languages. it can be a translated book or an internet page (wikipedia, news site) our algorithm tries to match Bashkir sentences with their translation in Russian We give these pairs to people to check @inproceedings{ title={Bashkir-Russian parallel corpora}, author={Iskander Shakirov, Aigiz Kunafin}, year={2023} } texttranslation1M<n<10M16 likes84 downloads2y agoHugging Face16DCAgent /bash_textbook_tasks_tracestext1K<n<10K0 likes81 downloads11mo agoHugging Face17DCAgent /bash_textbook_taskstext10K<n<100K0 likes80 downloads10mo agoHugging Face18failed09 /bashkir-wikipedia-monolingual Bashkir Wikipedia Monolingual Corpus Cleaned sentence-level Bashkir text from Wikipedia for pretraining, tokenizer training and linguistic research. Overview Sentence-level text extracted from the Bashkir Wikipedia dump (bawiki-20260801), cleaned and filtered with automated language identification. The cleaned configuration is the recommended default for language modelling, tokenization and linguistic research; precleaned is an earlier, lighter extraction kept… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-monolingual.texttext-generation1M<n<10M0 likes71 downloads7d agoHugging Face19DCAgent2 /dcagent2-terminal-bench-2-dcagent-bash-textbook-tasks-traces-20251122-123745textn<1K0 likes65 downloads10mo agoHugging Face20DCAgent /exp_rpt_stack-bash-withtests-gpt5mini_glm_4.7_traces_jupitertext10K<n<100K0 likes60 downloads6mo agoHugging Face21AISafety-Student /labeled-bashBench LLM Misbehavior Activation Dataset Dataset of labeled agent trajectory steps for use with steering vector / activation extraction. Source This dataset labels the trajectories found in mandliya/basharena-synthetic-trajectories. Structure Each row is ONE specific step or flagged action from the full original agent trajectory. Field Description id Unique entry UUID task_id Original BashArena task_id source_file Path to the original trajectory file… See the full description on the dataset page: https://huggingface.co/datasets/AISafety-Student/labeled-bashBench.tabulartext-classification1K<n<10K1 likes55 downloads6mo agoHugging Face22laion /dev_set_v2_a3_rl_laion_exp_rpt_stack_bash_v3_70_8B_20260825_134028text1K<n<10K0 likes54 downloads29d agoHugging Face23DCAgent /bash_textbook_tasks_glm_4.7_traces_jupitertext10K<n<100K0 likes51 downloads6mo agoHugging Face24DCAgent /exp_rpt_nemotron-bash-withtests-gpt5mini_glm_4.7_traces_jupitertext10K<n<100K0 likes50 downloads6mo agoHugging Face25DCAgent /rl__40GPU_base_32b__exp_rpt_nemotron-bash__sft_GLM-4-7-swesmithtext10K<n<100K0 likes47 downloads7mo agoHugging Face26open-athena /stack-bash-v3-qwen3.5-122b-32k-tracestext10K<n<100K0 likes47 downloads3mo agoHugging Face27DCAgent2 /terminal_bench_2_rl_base_exp_rpt_stack_bash_with_gpt5_90_20260223_182659textn<1K0 likes46 downloads7mo agoHugging Face28DCAgent2 /terminal_bench_2_rl_base_exp_rpt_stack_bash_90_20260223_182701textn<1K0 likes44 downloads7mo agoHugging Face29BashkirNLPWorld /bashkir-wiki-corpusgated Dataset Card for Bashkir Wikipedia Corpus Dataset Details Dataset Description The Bashkir Wikipedia Corpus is a collection of 43,926 articles from Bashkir Wikipedia and Wikibooks, totaling approximately 10.6 million tokens and 8.9 million words. The data has been extracted from official Wikimedia dumps and processed to provide clean, well‑structured text suitable for NLP tasks. The corpus includes article titles, full content, categories, source… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-wiki-corpus.texttext-generation10K<n<100K0 likes44 downloads27d agoHugging Face30DCAgent /exp_rpt_nemotron-bash-withtests_glm_4.7_traces_jupitertext10K<n<100K0 likes44 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.