arm
wav2vec2-large-xls-r-300m-armenianwhisper-large-v3-turbo-armenianArmoRM-Llama3-8B-v0.1yolov8n_handwritten_text_detectionQwen3.8-27B-Fable-Distill-Heretic-ara-GGUFCaptain-Eris_Violet-V0.420-12B-GGUF-ARM-ImatrixMN-BackyardAI-Party-12B-v1-GGUF-IQ-ARM-ImatrixCaptain-Eris-Diogenes_Twilight-V0.420-12B-GGUF-ARM-Imatrix
Datasets
All datasets matching “arm”the-pile-splitted
Dataset description
The pile is an 800GB dataset of english text
designed by EleutherAI to train large-scale language models. The original version of
the dataset can be found here.
The dataset is divided into 22 smaller high-quality datasets. For more information
each of them, please refer to the datasheet for the pile.
However, the current version of the dataset, available on the Hub, is not splitted accordingly.
We had to solve this problem in order to improve the user… See the full description on the dataset page: https://huggingface.co/datasets/ArmelR/the-pile-splitted.fafb-em-blocksCircuitSense
CircuitSense
This dataset is a comprehensive multimodal circuit question-answering benchmark designed to evaluate visual reasoning and problem-solving capabilities across three main domains: Perception, Analysis, and Design. The dataset contains structured question-answer pairs with accompanying visual content, targeting different engineering cognitive levels and reasoning tasks.
Dataset Structure
The dataset is organized into three primary folders, each containing… See the full description on the dataset page: https://huggingface.co/datasets/armanakbari4/CircuitSense.banana-vidorev3-synthetic-arms
Banana ViDoRe v3 Synthetic Arms
Domain-separated ViDoRe v3 synthetic training arms for finance and industrial adaptation.
The Hub dataset uses finance and industrial as dataset configs/subsets. Within each config, splits separate
vlm_in_batch, vlm_ocr_bm25, banana_fullpipe, and hybrid_vlm_ocr_bm25_banana_fullpipe.
Generated at: 2026-06-29T11:49:45.670912+00:00
Total JSONL rows across configs/splits: 151691.
Images are stored once per subset under… See the full description on the dataset page: https://huggingface.co/datasets/vkehfdl1/banana-vidorev3-synthetic-arms.skilltrainbench-public
skilltrainbench training tasks
The training half of the skilltrainbench benchmark suite: for each of the
four datasets, the dev_task_names of its pinned train/test split, in Harbor
task format.
The held-out/test tasks are not in this repository. Neither are the
published splits that are not the pin, nor the tasks that fall outside each
pinned split's task set. Use this repository for skill formation and
training; evaluate on the held-out half, which stays in the private source… See the full description on the dataset page: https://huggingface.co/datasets/armin-aptura/skilltrainbench-public.claude-fable-5-claude-code
claude-fable-5 Agent Traces
It's worth noting that our team was working with Glint-Research to collect as much fable data as possible.
These are just the anonymized raw traces of both of our teams combined. This means that Glint-Research/Fable-5-traces was created from formatting and splitting up this same dataset. If you use one for your tune, don't use the other (it's the same exact data).
For training on this dataset I recommend using the teich package to convert to openai… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/claude-fable-5-claude-code.

