CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yatin-superintelligence /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M52 likes5.9k downloads7mo agoHugging Face02Edge0 /gpa-demosaudion<1K12 likes783 downloads9mo agoHugging Face03BlueIsGreen /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/BlueIsGreen/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M11 likes620 downloads7mo agoHugging Face04DEMIRUNC /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/DEMIRUNC/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes551 downloads6mo agoHugging Face05edgeimpulse /Hey-Edge hey_edge — Wake Word Synthetic Speech Dataset Synthetic, augmented audio for training a small wake-word / keyword-spotting model. Generated with piper_tts and local audio augmentation. Classes Label Samples background_noise 200 hey_edge 378 unknown 1071 hey_edge — the target wake phrase and close variants. unknown — near-miss and unrelated short phrases. background_noise — synthetic background noise. Audio Specification… See the full description on the dataset page: https://huggingface.co/datasets/edgeimpulse/Hey-Edge.audioaudio-classification1K<n<10K0 likes516 downloads3mo agoHugging Face06Torenn /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/Torenn/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M1 likes434 downloads7mo agoHugging Face07ppenner /edge-agent-reasoning-websearch-260k Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/ppenner/edge-agent-reasoning-websearch-260k.texttext-generation100K<n<1M0 likes338 downloads4mo agoHugging Face08JACKYS999 /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/JACKYS999/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes269 downloads5mo agoHugging Face09svryn /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/svryn/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes184 downloads5mo agoHugging Face10Dorjzodovsuren /synthetic_mbspeech_dataset_edgettsaudio1K<n<10K0 likes22 downloads8mo agoHugging Face11szzs1693 /edge-ai-cough-count Edge Artificial Intelligence (edge-AI) Cough Counting This is a mirrored dataset of edge-AI Cough Counting Dataset, converted to Parquet format and hosted on Hugging Face for more accessibility. Additional preprocessing has been made, such as data aggregation and audio visualizations, for in-depth data details without downloading the whole dataset. Background Counting the number of times a patient coughs per day is an essential biomarker in determining treatment efficacy… See the full description on the dataset page: https://huggingface.co/datasets/szzs1693/edge-ai-cough-count.audioaudio-classification1K<n<10K0 likes19 downloads9mo agoHugging Face12CortexSwarm /EdgeMMEval EdgeMMEval Minimal multimodal evaluation dataset for on-device inference testing. Covers functional correctness, accuracy, latency stress, and memory pressure across image, audio, text, multi-turn, combination, structured output, and tool-calling cases. Dataset summary The test split is defined in data/test/metadata.jsonl (200 rows). Each row has a test_id (for example IMG-001, STO-020) and a modality. Modality Samples Focus Image 34 VQA, OCR, description… See the full description on the dataset page: https://huggingface.co/datasets/CortexSwarm/EdgeMMEval.audiovisual-question-answeringn<1K0 likes19 downloads5mo agoHugging Face13ollama456 /librispeech-whisper-edgeaudio1K<n<10K0 likes16 downloads7mo agoHugging Face14SheepHuan /EdgeEnergyBenchaudion<1K0 likes8 downloads4mo agoHugging Face15ggix /vl8-edge-noise-bank vl8-edge noise bank Curated noise bundle for vl8-edge-sdk's synth.augment pipeline (issue #9). Repackaged from the DEMAND corpus on Zenodo. License: CC-BY-4.0 Source: DEMAND — Joachim Thiemann, Nobutaka Ito, Emmanuel Vincent Format: 16 kHz mono 16-bit PCM, 6 channels (3 cafe + 3 street), ~40 MB The SDK lazy-downloads noise-bank-v1.tar.gz from this repo on first vl8 build --augment and verifies per-file SHA256 against the manifest at… See the full description on the dataset page: https://huggingface.co/datasets/ggix/vl8-edge-noise-bank.audion<1K0 likes2 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.