CoolFace
21 results

arm

ArmelR /the-pile-splitted Dataset description The pile is an 800GB dataset of english text designed by EleutherAI to train large-scale language models. The original version of the dataset can be found here. The dataset is divided into 22 smaller high-quality datasets. For more information each of them, please refer to the datasheet for the pile. However, the current version of the dataset, available on the Hub, is not splitted accordingly. We had to solve this problem in order to improve the user… See the full description on the dataset page: https://huggingface.co/datasets/ArmelR/the-pile-splitted.text10M<n<100M23 likes15k downloads3y agoHugging FaceArminshfard /fafb-em-blocks2 likes9.4k downloads3mo agoHugging Facearmanakbari4 /CircuitSense CircuitSense This dataset is a comprehensive multimodal circuit question-answering benchmark designed to evaluate visual reasoning and problem-solving capabilities across three main domains: Perception, Analysis, and Design. The dataset contains structured question-answer pairs with accompanying visual content, targeting different engineering cognitive levels and reasoning tasks. Dataset Structure The dataset is organized into three primary folders, each containing… See the full description on the dataset page: https://huggingface.co/datasets/armanakbari4/CircuitSense.imagequestion-answering1K<n<10K0 likes9.3k downloads1y agoHugging Facevkehfdl1 /banana-vidorev3-synthetic-arms Banana ViDoRe v3 Synthetic Arms Domain-separated ViDoRe v3 synthetic training arms for finance and industrial adaptation. The Hub dataset uses finance and industrial as dataset configs/subsets. Within each config, splits separate vlm_in_batch, vlm_ocr_bm25, banana_fullpipe, and hybrid_vlm_ocr_bm25_banana_fullpipe. Generated at: 2026-06-29T11:49:45.670912+00:00 Total JSONL rows across configs/splits: 151691. Images are stored once per subset under… See the full description on the dataset page: https://huggingface.co/datasets/vkehfdl1/banana-vidorev3-synthetic-arms.imagevisual-document-retrieval10K<n<100K0 likes7.7k downloads2mo agoHugging Facearmin-aptura /skilltrainbench-public skilltrainbench training tasks The training half of the skilltrainbench benchmark suite: for each of the four datasets, the dev_task_names of its pinned train/test split, in Harbor task format. The held-out/test tasks are not in this repository. Neither are the published splits that are not the pin, nor the tasks that fall outside each pinned split's task set. Use this repository for skill formation and training; evaluate on the held-out half, which stays in the private source… See the full description on the dataset page: https://huggingface.co/datasets/armin-aptura/skilltrainbench-public.other0 likes6.6k downloads3d agoHugging Facearmand0e /claude-fable-5-claude-code claude-fable-5 Agent Traces It's worth noting that our team was working with Glint-Research to collect as much fable data as possible. These are just the anonymized raw traces of both of our teams combined. This means that Glint-Research/Fable-5-traces was created from formatting and splitting up this same dataset. If you use one for your tune, don't use the other (it's the same exact data). For training on this dataset I recommend using the teich package to convert to openai… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/claude-fable-5-claude-code.tabulartext-generationn<1K387 likes4.9k downloads12d agoHugging Face