CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlx-community /optiq-code-traces OptiQ Code Traces Gold-verified agentic software-engineering trajectories, produced by OptiQ Code, the terminal coding agent for local models on a Mac. Each trajectory is a full tool-calling run against a real repository bug, and every resolved label is set by executing the gold tests (FAIL_TO_PASS + PASS_TO_PASS) after applying the model's patch, never by the agent's own self-report. The dataset is 1,789 agent sessions in HuggingFace Session-Traces format (the agent-traces… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/optiq-code-traces.tabulartext-generation1K<n<10K4 likes801 downloads9d agoHugging Face02mlx-community /ToolMind ToolMind: A Large-Scale, Reasoning-Enhanced Tool-Use Dataset ToolMind is a large-scale, high-quality tool-agentic dataset with 160k synthetic data instances generated using over 20k tools and 200k augmented open-source data instances. Our data synthesis pipeline first constructs a function graph based on parameter correlations and then uses a multi-agent framework to simulate realistic user–assistant–tool interactions. Beyond trajectory-level validation, we employ fine-grained… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/ToolMind.documenttext-generation100K<n<1M2 likes151 downloads4mo agoHugging Face03mlx-community /Apertus-v1.5-QAT-10K mlx-community/Apertus-v1.5-QAT-10K This is a 2000 sample subset of the chosen pairs inside swiss-ai/Apertus_v1p5_Preference_Data for MLX-LM-LoRA and MLX-LoRA-Studio and the Quantization Aware Trained Appertus models. texttext-generation10K<n<100K1 likes120 downloads7d agoHugging Face04mlx-community /JOSIE-v2-Instruct-5K JOSIE v2 Instruct 5K A high-quality instruction-following dataset featuring J.O.S.I.E. (Just One Super Intelligent Entity) - an advanced AI assistant with a distinctive personality combining intellectual rigor, dry wit, and genuine helpfulness. Dataset Overview Size: 5,000 conversational samples Format: JSONL (JSON Lines) Source Model: GPT-5.4-nano via OpenAI Batch API Use Case: Finetuning language models on Apple Silicon using mlx-lm or mlx-lm-lora License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/JOSIE-v2-Instruct-5K.texttext-generation1K<n<10K5 likes78 downloads5mo agoHugging Face05Goekdeniz-Guelmez /MLX-Benchmark-V2 MLX Benchmark Dataset Dataset Summary The MLX Benchmark Dataset is a curated evaluation benchmark consisting of 520 questions designed to measure large language model (LLM) proficiency in Apple's MLX machine learning framework. MLX is an array framework for machine learning on Apple Silicon that leverages unified memory architecture, and this dataset is the first comprehensive benchmark specifically targeting MLX knowledge and coding ability. The dataset covers the… See the full description on the dataset page: https://huggingface.co/datasets/Goekdeniz-Guelmez/MLX-Benchmark-V2.textquestion-answeringn<1K2 likes44 downloads5mo agoHugging Face06mlx-community /medfit-dataset MEDFIT Medical QA Dataset This dataset contains 6,444 unique healthcare-related question-answer pairs designed for fine-tuning language models for medical chatbot applications. The dataset was specifically curated for the MEDFIT-LLM research project focusing on domain-focused fine-tuning of small language models for healthcare applications. All credits for the methodology and dataset creation go to Aditya Karnam Gururaj Rao, Arjun Jaggi, and Sonam Naidu. The dataset has been… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/medfit-dataset.textquestion-answering1K<n<10K3 likes35 downloads1y agoHugging Face07mlx-community /JOSIE-DPO-Chosen-Ministral JOSIE-DPO-Chosen — Ministral Half of a DPO dataset. The chosen responses are here. You generate the rejected ones — and that's the point. Overview This dataset contains the chosen-only side of a preference dataset designed to align any LLM with the personality, tone, and response style of J.O.S.I.E. (Just One Super Intelligent Entity) — the viral model family created by Gökdeniz Gülmez. The chosen responses were generated by a fine-tuned Ministral-14B model… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/JOSIE-DPO-Chosen-Ministral.texttext-generation1K<n<10K1 likes31 downloads7mo agoHugging Face08sebastianboehler /hallmark-mlx-reviewed-policy-traces hallmark-mlx-reviewed-policy-traces Reviewed citation-verification training traces for hallmark-mlx. Contents train.jsonl: 75 supervised examples valid.jsonl: 6 supervised examples Source reviewed traces: reviewed_seed_traces_combined.jsonl with 45 full traces. Format Each row is a prepared supervised training example for MLX LoRA fine-tuning. The format is the exact snapshot used by the kept Qwen 1.5B run. Upload Note Review the… See the full description on the dataset page: https://huggingface.co/datasets/sebastianboehler/hallmark-mlx-reviewed-policy-traces.texttext-generationn<1K0 likes30 downloads6mo agoHugging Face09mlx-community /tnc-archive Paraacademic institution's archived educational activities metadata scraped and packed as a dataset. Dataset Details Dataset Description Scraped titles and summarized descriptions of the "non-members available" data of the Seminars of The New Centre for Research & Practice, took this from our website where i have a status of god of FireStoreStoNe (FSSN). Regarding the latter, should i not scrap the data with an access to the direct descriptions links? 100%… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/tnc-archive.textfeature-extractionn<1K1 likes25 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.