datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bashkir-frequency-index
Bashkir Frequency Index v11.5
Word-frequency index for Bashkir, computed over a large monolingual
Bashkir-language dataset, for NLP, spellchecking and lexical research.
Overview
Word-frequency index for the Bashkir language computed over a large monolingual
Bashkir-language dataset. Non-Bashkir admixture, borrowed vocabulary
and scanning artifacts were reduced with automated language filtering. The
public configuration (count ≥ 3) is the recommended default;… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-frequency-index.bashkir-wikipedia-parallel
Bashkir-Russian Wikipedia Parallel Corpus
Sentence-level Bashkir-Russian parallel text from Wikipedia, scored and filtered
for machine translation.
Overview
Sentence-level Bashkir-Russian parallel dataset extracted from the corresponding
Bashkir and Russian Wikipedia dumps dated 2026-08-01. Candidate pairs are scored
for semantic alignment with multilingual sentence encoders (Meta LASER3, Google
LaBSE) and the in-domain Bashkir-Russian Pair Scorer. The filtered… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-parallel.bashkir-ngram-index
Bashkir Word N-gram Index v11.5
Exact within-sentence word n-gram counts for Bashkir: unigrams, bigrams and
trigrams for spellchecking, OCR post-processing and lightweight language modelling.
Overview
Exact word n-gram counts derived from a monolingual Bashkir-language dataset. The
release provides unigram, bigram and trigram indexes for corpus processing,
spellchecking, OCR post-processing, autocomplete and lightweight language-model
experiments. The unigrams… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-ngram-index.uav-fault-symptom-reports
UAV Fault Symptom Reports
A synthetic dataset of UAV flight telemetry paired with operator-style symptom reports written by a
language model. Each row is one five-second window of a flight: 20 telemetry channels, the fault
class, a severity derived from simulated consequences, and a one-sentence report.
split
rows
flights
model-written reports
unique reports
benchmark
10,500
2,100
82.0%
79.7%
challenge
3,500
700
88.1%
91.0%
benchmark is balanced across seven… See the full description on the dataset page: https://huggingface.co/datasets/Bashifu/uav-fault-symptom-reports.labeled-bashBench
LLM Misbehavior Activation Dataset
Dataset of labeled agent trajectory steps for use with steering vector / activation extraction.
Source
This dataset labels the trajectories found in mandliya/basharena-synthetic-trajectories.
Structure
Each row is ONE specific step or flagged action from the full original agent trajectory.
Field
Description
id
Unique entry UUID
task_id
Original BashArena task_id
source_file
Path to the original trajectory file… See the full description on the dataset page: https://huggingface.co/datasets/AISafety-Student/labeled-bashBench.bimanual_so100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_aloha",
"total_episodes": 30,
"total_frames": 22836,
"total_tasks": 1,
"total_videos": 90,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bashar-Alhaffar/bimanual_so100.basharena-monitor-evalbasharena_action_only_sonnet45_largebashkir-news-binary
Dataset Card for Bashkir News Binary Classification Dataset
Dataset Details
Dataset Description
This dataset contains 16,994 Bashkir-language news and analytical articles labeled for binary classification: news (label=1) vs analytics (label=0). The dataset is perfectly balanced with 8,497 examples in each class. It was created to support NLP research and applications for the Bashkir language, a low-resource Turkic language.
Curated by: Arabov… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-news-binary.basharena_action_only_opus46_largeso100_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_aloha",
"total_episodes": 3,
"total_frames": 2024,
"total_tasks": 1,
"total_videos": 9,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bashar-Alhaffar/so100_test.bashkir-news-multilabel
Dataset Card for Bashkir News Multilabel Classification Dataset
Dataset Details
Dataset Description
This dataset contains 22,318 Bashkir-language news and analytical articles annotated with 14 thematic labels for multi-label text classification tasks. Each article can belong to several categories simultaneously. The average number of labels per article is 3.6. The dataset is designed to support NLP research and applications for the Bashkir language… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-news-multilabel.basharena_aware50k_nepali_chatbot_datasettitanic_datasetbasharena_xml_with_assistant_textbasharena_action_onlynepali_chatbot_datasetbasharena_action_only_xml_without_assistant_text
