datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ShIO-bash-26.1
ShIO-bash-26.1
Shell input-output (ShIO) Bash dataset produced by ShIOEnv, a Gymnasium-compatible Bash environment designed to collect execution-annotated command interactions in a Linux system.
Dataset summary
The dataset consists of command-line inputs paired with their execution artifacts, including observable outputs and a structured representation of environment state changes. Samples are produced by executing synthesized Bash inputs inside a… See the full description on the dataset page: https://huggingface.co/datasets/jragsdale1/ShIO-bash-26.1.bashkir-frequency-index
Bashkir Frequency Index v11.5
Word-frequency index for Bashkir, computed over a large monolingual
Bashkir-language dataset, for NLP, spellchecking and lexical research.
Overview
Word-frequency index for the Bashkir language computed over a large monolingual
Bashkir-language dataset. Non-Bashkir admixture, borrowed vocabulary
and scanning artifacts were reduced with automated language filtering. The
public configuration (count ≥ 3) is the recommended default;… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-frequency-index.bashkir-wikipedia-parallel
Bashkir-Russian Wikipedia Parallel Corpus
Sentence-level Bashkir-Russian parallel text from Wikipedia, scored and filtered
for machine translation.
Overview
Sentence-level Bashkir-Russian parallel dataset extracted from the corresponding
Bashkir and Russian Wikipedia dumps dated 2026-08-01. Candidate pairs are scored
for semantic alignment with multilingual sentence encoders (Meta LASER3, Google
LaBSE) and the in-domain Bashkir-Russian Pair Scorer. The filtered… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-parallel.bashkir-ngram-index
Bashkir Word N-gram Index v11.5
Exact within-sentence word n-gram counts for Bashkir: unigrams, bigrams and
trigrams for spellchecking, OCR post-processing and lightweight language modelling.
Overview
Exact word n-gram counts derived from a monolingual Bashkir-language dataset. The
release provides unigram, bigram and trigram indexes for corpus processing,
spellchecking, OCR post-processing, autocomplete and lightweight language-model
experiments. The unigrams… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-ngram-index.bimanual_so100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_aloha",
"total_episodes": 30,
"total_frames": 22836,
"total_tasks": 1,
"total_videos": 90,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bashar-Alhaffar/bimanual_so100.Bwenge
Dataset Description
This dataset was created to develop a machine translation model for bidirectional translation between Kinyarwanda and English for education-based sentences, in particular for the Atingi learning platform.
Repository:link to the GitHub repository containing the code for training the model on this data, and the code for the collection of the monolingual data.
Data Format: TSV
Model: huggingface model link.
Dataset Summary
Data… See the full description on the dataset page: https://huggingface.co/datasets/Coding-With-Bashir/Bwenge.labeled-bashBench
LLM Misbehavior Activation Dataset
Dataset of labeled agent trajectory steps for use with steering vector / activation extraction.
Source
This dataset labels the trajectories found in mandliya/basharena-synthetic-trajectories.
Structure
Each row is ONE specific step or flagged action from the full original agent trajectory.
Field
Description
id
Unique entry UUID
task_id
Original BashArena task_id
source_file
Path to the original trajectory file… See the full description on the dataset page: https://huggingface.co/datasets/AISafety-Student/labeled-bashBench.rlvr-bash-terminal-bench
rlvr-bash-terminal-bench
RLVR (Reinforcement Learning with Verifiable Rewards) dataset for bash scripting, generated from Terminal-Bench tasks.
Stats
Metric
Value
Total samples
1,120
Unique tasks
88
Avg samples/task
12.7
Average reward
0.249
Perfect solutions (reward=1.0)
10.4%
Partial solutions (0<reward<1)
28.8%
Zero reward
60.8%
Tasks fully solved
13.6%
Format
{
"task_id": "string",
"prompt": "string",
"completion":… See the full description on the dataset page: https://huggingface.co/datasets/lisayan/rlvr-bash-terminal-bench.bashkir-russian-parallel
Dataset Card for Bashkir-Russian Parallel Corpus
Dataset Details
Dataset Description
Bashkir-Russian Parallel Corpus is a large-scale sentence-aligned parallel corpus for the Bashkir–Russian language pair, assembled from authentic human-created translations. It contains 3,040,085 unique parallel sentence pairs, where each Bashkir sentence is aligned with its Russian counterpart.
The corpus combines data from three open parallel corpora: TIL-MT… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-russian-parallel.bashkir-news-binary
Dataset Card for Bashkir News Binary Classification Dataset
Dataset Details
Dataset Description
This dataset contains 16,994 Bashkir-language news and analytical articles labeled for binary classification: news (label=1) vs analytics (label=0). The dataset is perfectly balanced with 8,497 examples in each class. It was created to support NLP research and applications for the Bashkir language, a low-resource Turkic language.
Curated by: Arabov… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-news-binary.basharena_action_only_sonnet45_largebasharena-monitor-evalbasharena_action_only_opus46_largeso100_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_aloha",
"total_episodes": 3,
"total_frames": 2024,
"total_tasks": 1,
"total_videos": 9,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bashar-Alhaffar/so100_test.bashkir-news-multilabel
Dataset Card for Bashkir News Multilabel Classification Dataset
Dataset Details
Dataset Description
This dataset contains 22,318 Bashkir-language news and analytical articles annotated with 14 thematic labels for multi-label text classification tasks. Each article can belong to several categories simultaneously. The average number of labels per article is 3.6. The dataset is designed to support NLP research and applications for the Bashkir language… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-news-multilabel.css-colorstitanic_datasetbash-toolcallsbasharena_aware50k_nepali_chatbot_datasetbasharena_action_onlynepali_chatbot_datasetbasharena_xml_with_assistant_textbasharena_action_only_xml_without_assistant_textEDA_Assignment
🎓 Student Dropout Prediction Dataset — EDA Assignment
By Tomer Bash | Data Science Course — Assignment #1
📹 Presentation Video
Presentation Video link - https://youtu.be/KyafBx9W7Qg
📌 Dataset Overview
Property
Details
Source
Kaggle
Rows
4,424 students
Features
35 columns
Target Variable
Target — Graduate, Enrolled, Dropout
Task Type
Multi-class Classification
The dataset contains demographic, financial, academic, and… See the full description on the dataset page: https://huggingface.co/datasets/Bashifu/EDA_Assignment.testtourism-package-prediction
