CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01failed09 /bashkir-frequency-index Bashkir Frequency Index v11.5 Word-frequency index for Bashkir, computed over a large monolingual Bashkir-language dataset, for NLP, spellchecking and lexical research. Overview Word-frequency index for the Bashkir language computed over a large monolingual Bashkir-language dataset. Non-Bashkir admixture, borrowed vocabulary and scanning artifacts were reduced with automated language filtering. The public configuration (count ≥ 3) is the recommended default;… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-frequency-index.tabulartext-classification1M<n<10M0 likes238 downloads3d agoHugging Face02failed09 /bashkir-wikipedia-parallel Bashkir-Russian Wikipedia Parallel Corpus Sentence-level Bashkir-Russian parallel text from Wikipedia, scored and filtered for machine translation. Overview Sentence-level Bashkir-Russian parallel dataset extracted from the corresponding Bashkir and Russian Wikipedia dumps dated 2026-08-01. Candidate pairs are scored for semantic alignment with multilingual sentence encoders (Meta LASER3, Google LaBSE) and the in-domain Bashkir-Russian Pair Scorer. The filtered… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-parallel.tabulartranslation100K<n<1M0 likes162 downloads7d agoHugging Face03failed09 /bashkir-ngram-index Bashkir Word N-gram Index v11.5 Exact within-sentence word n-gram counts for Bashkir: unigrams, bigrams and trigrams for spellchecking, OCR post-processing and lightweight language modelling. Overview Exact word n-gram counts derived from a monolingual Bashkir-language dataset. The release provides unigram, bigram and trigram indexes for corpus processing, spellchecking, OCR post-processing, autocomplete and lightweight language-model experiments. The unigrams… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-ngram-index.tabulartext-classification10M<n<100M0 likes104 downloads3d agoHugging Face04Bashifu /uav-fault-symptom-reports UAV Fault Symptom Reports A synthetic dataset of UAV flight telemetry paired with operator-style symptom reports written by a language model. Each row is one five-second window of a flight: 20 telemetry channels, the fault class, a severity derived from simulated consequences, and a one-sentence report. split rows flights model-written reports unique reports benchmark 10,500 2,100 82.0% 79.7% challenge 3,500 700 88.1% 91.0% benchmark is balanced across seven… See the full description on the dataset page: https://huggingface.co/datasets/Bashifu/uav-fault-symptom-reports.imagetext-classification10K<n<100K0 likes59 downloads19h agoHugging Face05AISafety-Student /labeled-bashBench LLM Misbehavior Activation Dataset Dataset of labeled agent trajectory steps for use with steering vector / activation extraction. Source This dataset labels the trajectories found in mandliya/basharena-synthetic-trajectories. Structure Each row is ONE specific step or flagged action from the full original agent trajectory. Field Description id Unique entry UUID task_id Original BashArena task_id source_file Path to the original trajectory file… See the full description on the dataset page: https://huggingface.co/datasets/AISafety-Student/labeled-bashBench.tabulartext-classification1K<n<10K1 likes55 downloads6mo agoHugging Face06Bashar-Alhaffar /bimanual_so100This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100_aloha", "total_episodes": 30, "total_frames": 22836, "total_tasks": 1, "total_videos": 90, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:30" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bashar-Alhaffar/bimanual_so100.tabularrobotics10K<n<100K1 likes41 downloads1y agoHugging Face07abhayesian /basharena-monitor-evaltabularn<1K0 likes28 downloads6mo agoHugging Face08adityaasinha28 /basharena_action_only_sonnet45_largetabular1K<n<10K0 likes28 downloads6mo agoHugging Face09BashkirNLPWorld /bashkir-news-binarygated Dataset Card for Bashkir News Binary Classification Dataset Dataset Details Dataset Description This dataset contains 16,994 Bashkir-language news and analytical articles labeled for binary classification: news (label=1) vs analytics (label=0). The dataset is perfectly balanced with 8,497 examples in each class. It was created to support NLP research and applications for the Bashkir language, a low-resource Turkic language. Curated by: Arabov… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-news-binary.tabulartext-classification10K<n<100K0 likes26 downloads27d agoHugging Face10adityaasinha28 /basharena_action_only_opus46_largetabular1K<n<10K0 likes25 downloads6mo agoHugging Face11Bashar-Alhaffar /so100_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100_aloha", "total_episodes": 3, "total_frames": 2024, "total_tasks": 1, "total_videos": 9, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:3" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bashar-Alhaffar/so100_test.tabularrobotics1K<n<10K0 likes21 downloads1y agoHugging Face12BashkirNLPWorld /bashkir-news-multilabelgated Dataset Card for Bashkir News Multilabel Classification Dataset Dataset Details Dataset Description This dataset contains 22,318 Bashkir-language news and analytical articles annotated with 14 thematic labels for multi-label text classification tasks. Each article can belong to several categories simultaneously. The average number of labels per article is 3.6. The dataset is designed to support NLP research and applications for the Bashkir language… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-news-multilabel.tabulartext-classification10K<n<100K0 likes20 downloads27d agoHugging Face13adityaasinha28 /basharena_awaretabularn<1K0 likes13 downloads7mo agoHugging Face14bashyaldhiraj2067 /50k_nepali_chatbot_datasettabular10K<n<100K0 likes13 downloads6mo agoHugging Face15bashyaldhiraj2067 /titanic_datasettabularn<1K0 likes13 downloads6mo agoHugging Face16adityaasinha28 /basharena_xml_with_assistant_texttabularn<1K0 likes11 downloads7mo agoHugging Face17adityaasinha28 /basharena_action_onlytabularn<1K0 likes11 downloads7mo agoHugging Face18bashyaldhiraj2067 /nepali_chatbot_datasettabular10K<n<100K0 likes11 downloads6mo agoHugging Face19adityaasinha28 /basharena_action_only_xml_without_assistant_texttabularn<1K0 likes8 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.