CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jragsdale1 /ShIO-bash-26.1 ShIO-bash-26.1 Shell input-output (ShIO) Bash dataset produced by ShIOEnv, a Gymnasium-compatible Bash environment designed to collect execution-annotated command interactions in a Linux system. Dataset summary The dataset consists of command-line inputs paired with their execution artifacts, including observable outputs and a structured representation of environment state changes. Samples are produced by executing synthesized Bash inputs inside a… See the full description on the dataset page: https://huggingface.co/datasets/jragsdale1/ShIO-bash-26.1.imagetext-generation1M<n<10M1 likes321 downloads4mo agoHugging Face02failed09 /bashkir-frequency-index Bashkir Frequency Index v11.5 Word-frequency index for Bashkir, computed over a large monolingual Bashkir-language dataset, for NLP, spellchecking and lexical research. Overview Word-frequency index for the Bashkir language computed over a large monolingual Bashkir-language dataset. Non-Bashkir admixture, borrowed vocabulary and scanning artifacts were reduced with automated language filtering. The public configuration (count ≥ 3) is the recommended default;… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-frequency-index.tabulartext-classification1M<n<10M0 likes219 downloads21h agoHugging Face03failed09 /bashkir-wikipedia-parallel Bashkir-Russian Wikipedia Parallel Corpus Sentence-level Bashkir-Russian parallel text from Wikipedia, scored and filtered for machine translation. Overview Sentence-level Bashkir-Russian parallel dataset extracted from the corresponding Bashkir and Russian Wikipedia dumps dated 2026-08-01. Candidate pairs are scored for semantic alignment with multilingual sentence encoders (Meta LASER3, Google LaBSE) and the in-domain Bashkir-Russian Pair Scorer. The filtered… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-parallel.tabulartranslation100K<n<1M0 likes180 downloads4d agoHugging Face04failed09 /bashkir-ngram-index Bashkir Word N-gram Index v11.5 Exact within-sentence word n-gram counts for Bashkir: unigrams, bigrams and trigrams for spellchecking, OCR post-processing and lightweight language modelling. Overview Exact word n-gram counts derived from a monolingual Bashkir-language dataset. The release provides unigram, bigram and trigram indexes for corpus processing, spellchecking, OCR post-processing, autocomplete and lightweight language-model experiments. The unigrams… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-ngram-index.tabulartext-classification10M<n<100M0 likes84 downloads21h agoHugging Face05Bashar-Alhaffar /bimanual_so100This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100_aloha", "total_episodes": 30, "total_frames": 22836, "total_tasks": 1, "total_videos": 90, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:30" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bashar-Alhaffar/bimanual_so100.tabularrobotics10K<n<100K1 likes41 downloads1y agoHugging Face06Coding-With-Bashir /Bwenge Dataset Description This dataset was created to develop a machine translation model for bidirectional translation between Kinyarwanda and English for education-based sentences, in particular for the Atingi learning platform. Repository:link to the GitHub repository containing the code for training the model on this data, and the code for the collection of the monolingual data. Data Format: TSV Model: huggingface model link. Dataset Summary Data… See the full description on the dataset page: https://huggingface.co/datasets/Coding-With-Bashir/Bwenge.tabulartranslation10K<n<100K0 likes40 downloads4mo agoHugging Face07AISafety-Student /labeled-bashBench LLM Misbehavior Activation Dataset Dataset of labeled agent trajectory steps for use with steering vector / activation extraction. Source This dataset labels the trajectories found in mandliya/basharena-synthetic-trajectories. Structure Each row is ONE specific step or flagged action from the full original agent trajectory. Field Description id Unique entry UUID task_id Original BashArena task_id source_file Path to the original trajectory file… See the full description on the dataset page: https://huggingface.co/datasets/AISafety-Student/labeled-bashBench.tabulartext-classification1K<n<10K1 likes38 downloads6mo agoHugging Face08lisayan /rlvr-bash-terminal-bench rlvr-bash-terminal-bench RLVR (Reinforcement Learning with Verifiable Rewards) dataset for bash scripting, generated from Terminal-Bench tasks. Stats Metric Value Total samples 1,120 Unique tasks 88 Avg samples/task 12.7 Average reward 0.249 Perfect solutions (reward=1.0) 10.4% Partial solutions (0<reward<1) 28.8% Zero reward 60.8% Tasks fully solved 13.6% Format { "task_id": "string", "prompt": "string", "completion":… See the full description on the dataset page: https://huggingface.co/datasets/lisayan/rlvr-bash-terminal-bench.tabulartext-generation1K<n<10K0 likes34 downloads8mo agoHugging Face09BashkirNLPWorld /bashkir-russian-parallelgated Dataset Card for Bashkir-Russian Parallel Corpus Dataset Details Dataset Description Bashkir-Russian Parallel Corpus is a large-scale sentence-aligned parallel corpus for the Bashkir–Russian language pair, assembled from authentic human-created translations. It contains 3,040,085 unique parallel sentence pairs, where each Bashkir sentence is aligned with its Russian counterpart. The corpus combines data from three open parallel corpora: TIL-MT… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-russian-parallel.tabulartext-generation1M<n<10M0 likes27 downloads5d agoHugging Face10BashkirNLPWorld /bashkir-news-binarygated Dataset Card for Bashkir News Binary Classification Dataset Dataset Details Dataset Description This dataset contains 16,994 Bashkir-language news and analytical articles labeled for binary classification: news (label=1) vs analytics (label=0). The dataset is perfectly balanced with 8,497 examples in each class. It was created to support NLP research and applications for the Bashkir language, a low-resource Turkic language. Curated by: Arabov… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-news-binary.tabulartext-classification10K<n<100K0 likes26 downloads25d agoHugging Face11adityaasinha28 /basharena_action_only_sonnet45_largetabular1K<n<10K0 likes26 downloads6mo agoHugging Face12abhayesian /basharena-monitor-evaltabularn<1K0 likes25 downloads6mo agoHugging Face13adityaasinha28 /basharena_action_only_opus46_largetabular1K<n<10K0 likes23 downloads6mo agoHugging Face14Bashar-Alhaffar /so100_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100_aloha", "total_episodes": 3, "total_frames": 2024, "total_tasks": 1, "total_videos": 9, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:3" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bashar-Alhaffar/so100_test.tabularrobotics1K<n<10K0 likes21 downloads1y agoHugging Face15BashkirNLPWorld /bashkir-news-multilabelgated Dataset Card for Bashkir News Multilabel Classification Dataset Dataset Details Dataset Description This dataset contains 22,318 Bashkir-language news and analytical articles annotated with 14 thematic labels for multi-label text classification tasks. Each article can belong to several categories simultaneously. The average number of labels per article is 3.6. The dataset is designed to support NLP research and applications for the Bashkir language… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-news-multilabel.tabulartext-classification10K<n<100K0 likes20 downloads25d agoHugging Face16Bashroom /css-colorstabularn<1K1 likes15 downloads2y agoHugging Face17bashyaldhiraj2067 /titanic_datasettabularn<1K0 likes15 downloads6mo agoHugging Face18kth8 /bash-toolcallstabularn<1K0 likes14 downloads4mo agoHugging Face19adityaasinha28 /basharena_awaretabularn<1K0 likes13 downloads7mo agoHugging Face20bashyaldhiraj2067 /50k_nepali_chatbot_datasettabular10K<n<100K0 likes13 downloads6mo agoHugging Face21adityaasinha28 /basharena_action_onlytabularn<1K0 likes11 downloads7mo agoHugging Face22bashyaldhiraj2067 /nepali_chatbot_datasettabular10K<n<100K0 likes11 downloads6mo agoHugging Face23adityaasinha28 /basharena_xml_with_assistant_texttabularn<1K0 likes10 downloads7mo agoHugging Face24adityaasinha28 /basharena_action_only_xml_without_assistant_texttabularn<1K0 likes8 downloads7mo agoHugging Face25Bashifu /EDA_Assignment 🎓 Student Dropout Prediction Dataset — EDA Assignment By Tomer Bash | Data Science Course — Assignment #1 📹 Presentation Video Presentation Video link - https://youtu.be/KyafBx9W7Qg 📌 Dataset Overview Property Details Source Kaggle Rows 4,424 students Features 35 columns Target Variable Target — Graduate, Enrolled, Dropout Task Type Multi-class Classification The dataset contains demographic, financial, academic, and… See the full description on the dataset page: https://huggingface.co/datasets/Bashifu/EDA_Assignment.tabular1K<n<10K0 likes4 downloads4mo agoHugging Face26BashyBaranaba /testtabularn<1K0 likes2 downloads4y agoHugging Face27Bash18 /tourism-package-predictiontabular1K<n<10K0 likes1 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.