CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ByteDance-Seed /EdgeBench Overview EdgeBench is a benchmark of 134 real-world tasks for evaluating how autonomous AI agents learn from real-world environments. Instead of measuring one-shot performance, EdgeBench places agents in executable task environments with realistic, multi-level feedback and lets them iterate for 12+ hours per task — tracking the full trajectory of improvement, not just the final score. We publicly release 51 tasks… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/EdgeBench.texttext-generationn<1K84 likes7.1k downloads2mo agoHugging Face02Edge0 /ark-asr-open-asr-leaderboard-results ARK-ASR Open ASR Leaderboard Results This dataset contains JSONL prediction manifests for AutoArk-AI/ARK-ASR-0.6B on hf-audio/open-asr-leaderboard public English short-form splits. These files are intended for Open ASR Leaderboard maintainer verification. Scoring summary from normalizer.eval_utils.score_results: Split WER RTFx ami/test 10.02 352.12 earnings22/test 9.77 331.88 gigaspeech/test 8.00 217.72 librispeech/test.clean 1.53 412.12 librispeech/test.other… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-open-asr-leaderboard-results.tabular10K<n<100K13 likes194 downloads3mo agoHugging Face03Edge0 /ark-asr-3b-open-asr-leaderboard-results ARK-ASR-3B Open ASR Leaderboard Results Raw JSONL manifests for AutoArk-AI/ARK-ASR-3B on the public English short-form hf-audio/open-asr-leaderboard splits. These manifests were generated on a local 8x RTX 4090 machine and scored with the shared Open ASR Leaderboard scorer: PYTHONPATH=. python - <<'PY' from normalizer.eval_utils import score_results score_results( 'ark_asr/results.AutoArk-AI-ARK-ASR-3B_20260622_official', 'AutoArk-AI/ARK-ASR-3B', ) PY Important:… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-3b-open-asr-leaderboard-results.tabularautomatic-speech-recognition10K<n<100K12 likes185 downloads3mo agoHugging Face04Edgerunners /Changelog-Nightly-Repositoriesarchive of all the repositories incl. metadata of: https://changelog.com/nightly will be used to train a spam classifier with spacy; hence the "text" column, but kept submeta in case this is useful for anyone else to re-format. The Dataset is provided ""AS IS"" and ""AS AVAILABLE"" without warranty of any kind, express or implied, including but not limited to warranties of merchantability, fitness for a particular purpose, title, or non-infringement. The Provider disclaims all liability for… See the full description on the dataset page: https://huggingface.co/datasets/Edgerunners/Changelog-Nightly-Repositories.text10K<n<100K0 likes158 downloads2y agoHugging Face05OniReimu /Edge-Computing-JEV EdgeIntent v1 EdgeIntent v1 is a benchmark of natural-language requests to edge services, each paired with the typed intent contract it expresses. It was built for the paper Replacing Large Language Models with Jev Decision Models for Low-Latency Edge Service Orchestration Delong Li, Xu Wang, Haochen Gong, Rui Lang, and Guangsheng Yu. University of Technology Sydney. Code, evaluation harness, and reproduction instructions: https://github.com/OniReimu/Edge-Computing-JEV This… See the full description on the dataset page: https://huggingface.co/datasets/OniReimu/Edge-Computing-JEV.texttext-classification10K<n<100K0 likes141 downloads1d agoHugging Face06ulamai /EdgeReason EdgeReason EdgeReason is a compact verifier-backed dataset for improving small and edge-deployable language models on tool use, structured JSON outputs, state/table/unit/date reasoning, compact Mathlib-derived SFT, and routing between direct answer, tool use, retrieval, clarification, and escalation. The dataset is designed for teams training small models with SFT, DPO, RLVR, GRPO, rejection sampling, and internal evaluation loops. It is not tied to any model vendor or… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/EdgeReason.texttext-generation10K<n<100K1 likes105 downloads3mo agoHugging Face07NeroSeungSan /synthengine-cot-edge-case-v1 SynthEngine CoT Edge Case Dataset v1.0 Premium synthetic Chain-of-Thought reasoning data for autonomous driving, robotics, and embodied AI edge cases. 🔗 Full dataset (1000 records) available on Gumroad This HuggingFace repo contains a free sample (10 records) under CC BY-NC-SA 4.0. 🎯 Why This Dataset? In 2025, NVIDIA Alpamayo-R1 proved that Chain-of-Causation reasoning improves autonomous driving planning accuracy by +12% and reduces close encounters by -35%.… See the full description on the dataset page: https://huggingface.co/datasets/NeroSeungSan/synthengine-cot-edge-case-v1.text10K<n<100K0 likes80 downloads4mo agoHugging Face08alirezaaminzadeh /slicenet-edge-slicing-scenarios SliceNet Edge Slicing Scenarios Synthetic edge computing and 5G network slicing placement scenarios for telecom, cloud, smart factory, connected vehicles, gaming, video, digital health, IoT, and smart city workloads. Samples File Industry Policy sample_smart_city.json Smart City / IoT Min Latency sample_connected_vehicle.json Automotive V2X Min Latency sample_online_gaming.json Cloud Gaming Min Latency sample_smart_factory.json Industry 4.0 Min… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/slicenet-edge-slicing-scenarios.textn<1K0 likes55 downloads2mo agoHugging Face09ssam17 /Edge-Industrial-Anomaly-Phi3 Edge-Industrial-Anomaly-Phi3: A Curated Dataset for SLMs This dataset is a curated collection of industrial sensor data formatted specifically for Small Language Models (SLMs) like Phi-3. It merges three high-value industrial domains into a unified "Natural Language Reasoning" format to move beyond simple binary classification. 🚀 Purpose Standard anomaly detection uses CSVs and Scikit-Learn. This dataset enables Generative Anomaly Detection, where a model like Phi-3 can… See the full description on the dataset page: https://huggingface.co/datasets/ssam17/Edge-Industrial-Anomaly-Phi3.texttext-generation10K<n<100K1 likes44 downloads9mo agoHugging Face10dougdotcon /douvras-network-edge-telemetry Douvras Network and Edge Telemetry v0.1 Synthetic TCP/UDP telemetry episodes with latency, packet loss, throughput, CPU and queue depth. Labels cover normal operation, latency, loss, congestion and resource saturation. It contains 60 records (40/10/10) across 12 episodes, split by episode. No real network traffic is included. This is a diagnostic benchmark only; it never remediates or changes a network. textn<1K0 likes34 downloads14d agoHugging Face11TCLResearchEurope /EdgeWisePersona Dataset Card for EdgeWisePersona Dataset Summary The core component of the dataset consists of natural language sessions between users and their smart home systems. These dialogues simulate realistic, free-form interactions in which users express commands, preferences, or queries. The sessions are grounded in underlying formalized behavioral routines. These routines, along with the user profiles they compose, are also included in the dataset as ground truth… See the full description on the dataset page: https://huggingface.co/datasets/TCLResearchEurope/EdgeWisePersona.texttext-generationn<1K5 likes32 downloads4mo agoHugging Face12edge2992 /github-issuestabular1K<n<10K0 likes30 downloads5y agoHugging Face13Edgerunners /NobodyExistsOnTheInternet_ToxicDPOqa_llama_factoryreformat of: NobodyExistsOnTheInternet/ToxicDPOqa for llama-factory DPO format usage example: "toxic_dpo_reformat": { "hf_hub_url": "lucyknada/NobodyExistsOnTheInternet_ToxicDPOqa_llama_factory", "ranking": true, "columns": { "prompt": "prompt", "response": "response", "system": "system" } }, text1K<n<10K0 likes30 downloads2y agoHugging Face14Edgerunners /Phoebus-86kraw human created erotica stories very aggresively cleaned version of Phoebus-127k to try to remove: product spambots, the website was being spammed a few times (usually with html tags) warning only pages ("Warning" and nothing else) edits (authors adding editorial history footnotes) patreon and alike callouts (author asking for donations) author notes, summaries, tagging some still got through, classifier used: Edgerunners/Phoebus-Spam-Classifier-v2 The Dataset is provided ""AS IS"" and… See the full description on the dataset page: https://huggingface.co/datasets/Edgerunners/Phoebus-86k.text10K<n<100K3 likes28 downloads2y agoHugging Face15EDGEwww25 /EDGE-DatasetThis is the dataset repository of paper EDGE: Enhanced Grounded GUI Understanding with Enriched Multi-Granularity Synthetic Data. Considering the huge number of images, the all_items.jsonl provided here contains the final QA pairs for training, but does not contain images for the time being. We will release all images as soon as possible. You can also follow the mark_webpages and dataset.py scripts provided in the code repository to generate your own webpage image-question-answering dataset.… See the full description on the dataset page: https://huggingface.co/datasets/EDGEwww25/EDGE-Dataset.textquestion-answering1M<n<10M0 likes28 downloads2y agoHugging Face16Inkwell-Software /screenplay-format-edge-cases Screenplay Format Edge Cases 48 original Fountain specimens in 24 contrast pairs — two near-identical inputs per pair, at the points where the Fountain syntax leaves a choice. In 17 pairs the one difference changes how the lines are classified. In the other 7 it changes the surface and the labels hold: a lowercase scene prefix, a cue extension, a non-Latin cue, escaped characters, a dual-dialogue caret, an inline note and centered-text markers. Version: 1.0.0 · Maintainer:… See the full description on the dataset page: https://huggingface.co/datasets/Inkwell-Software/screenplay-format-edge-cases.texttext-classificationn<1K1 likes25 downloads1d agoHugging Face17Edgerunners /glaive-first-human-only-dedupedglaive function calling dataset filtered just for the original human prompt and deduplicated afterwards The Dataset is provided ""AS IS"" and ""AS AVAILABLE"" without warranty of any kind, express or implied, including but not limited to warranties of merchantability, fitness for a particular purpose, title, or non-infringement. The Provider disclaims all liability for any damages or losses resulting from the use or misuse of the Dataset, including but not limited to any damages or losses… See the full description on the dataset page: https://huggingface.co/datasets/Edgerunners/glaive-first-human-only-deduped.text10K<n<100K0 likes19 downloads2y agoHugging Face18mukunda1729 /token-counting-edge-cases token-counting-edge-cases 20 short strings with approximate token counts across three tokenizer families: Claude, GPT (cl100k_base), and Llama (SentencePiece). Built for sanity-checking token counters / chunkers / context-window fitters. The numbers are approximate — exact counts depend on tokenizer version, BOS/EOS handling, and surrounding context. Expect ±1–2 token jitter. Use these to catch order-of-magnitude bugs (e.g. "your counter says 200 tokens for one emoji"), not as… See the full description on the dataset page: https://huggingface.co/datasets/mukunda1729/token-counting-edge-cases.textn<1K0 likes19 downloads5mo agoHugging Face19CortexSwarm /EdgeMMEval EdgeMMEval Minimal multimodal evaluation dataset for on-device inference testing. Covers functional correctness, accuracy, latency stress, and memory pressure across image, audio, text, multi-turn, combination, structured output, and tool-calling cases. Dataset summary The test split is defined in data/test/metadata.jsonl (200 rows). Each row has a test_id (for example IMG-001, STO-020) and a modality. Modality Samples Focus Image 34 VQA, OCR, description… See the full description on the dataset page: https://huggingface.co/datasets/CortexSwarm/EdgeMMEval.audiovisual-question-answeringn<1K0 likes19 downloads5mo agoHugging Face20Edgerunners /Phoebus-127k-labelsPhoebus-127k but with labels added so users can connect chapters together. raw human created erotica stories, needs filtering things that need to be filtered: product spambots, the website was being spammed a few times (usually with html tags) warning only pages ("Warning" and nothing else) edits (authors adding editorial history footnotes) patreon and alike callouts (author asking for donations) author notes, summaries, tagging The Dataset is provided ""AS IS"" and ""AS AVAILABLE""… See the full description on the dataset page: https://huggingface.co/datasets/Edgerunners/Phoebus-127k-labels.tabular100K<n<1M3 likes17 downloads2y agoHugging Face21Edgerunners /Phoebus-127kraw human created erotica stories, needs filtering things that need to be filtered: product spambots, the website was being spammed a few times (usually with html tags) warning only pages ("Warning" and nothing else) edits (authors adding editorial history footnotes) patreon and alike callouts (author asking for donations) author notes, summaries, tagging The Dataset is provided ""AS IS"" and ""AS AVAILABLE"" without warranty of any kind, express or implied, including but not limited to… See the full description on the dataset page: https://huggingface.co/datasets/Edgerunners/Phoebus-127k.text100K<n<1M3 likes16 downloads2y agoHugging Face22dvyomkesh /nemo-grpo-from083-full-edge-curation Nemotron 0.83 Edge-Prompt Curation This private dataset contains edge-prompt curation rollouts for the DGXChen/Tong CoT dataset. Seed edge prompts: 134 New rollout rows after seed exclusion: 7668 New edge prompts: 1648 Full edge prompts, seed plus rollout: 1782 Full dataset rows: 7830 Edge rate over full dataset: 0.2276 The Hugging Face dataset viewer is configured to load only data/full_edge_prompts_seed_plus_rollout.jsonl. The larger rollout and metadata files remain… See the full description on the dataset page: https://huggingface.co/datasets/dvyomkesh/nemo-grpo-from083-full-edge-curation.tabular1K<n<10K0 likes15 downloads4mo agoHugging Face23North-ML1 /wind-edge-1.6-sft Wind Lite SFT Custom supervised fine-tuning dataset for Wind Lite 1.6 by North AI. Dataset Summary 20,000 high-quality instruction-response pairs covering identity grounding, math reasoning, coding, general knowledge, and multi-turn conversations. Data Composition Category Count Description Math & Reasoning ~7,000 Arithmetic, algebra, percentages, unit conversions — with step-by-step working Coding ~4,000 Python, JavaScript, SQL, systems — with… See the full description on the dataset page: https://huggingface.co/datasets/North-ML1/wind-edge-1.6-sft.texttext-generation10K<n<100K0 likes14 downloads6mo agoHugging Face24open-llm-leaderboard /Edgerunners__meta-llama-3-8b-instruct-hf-ortho-baukit-34fail-3000total-bf16-detailsgated Dataset Card for Evaluation run of Edgerunners/meta-llama-3-8b-instruct-hf-ortho-baukit-34fail-3000total-bf16 Dataset automatically created during the evaluation run of model Edgerunners/meta-llama-3-8b-instruct-hf-ortho-baukit-34fail-3000total-bf16 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Edgerunners__meta-llama-3-8b-instruct-hf-ortho-baukit-34fail-3000total-bf16-details.tabular10K<n<100K0 likes13 downloads2y agoHugging Face25open-llm-leaderboard /DreadPoor__Blunt_Edge-8B-SLERP-detailsgated Dataset Card for Evaluation run of DreadPoor/Blunt_Edge-8B-SLERP Dataset automatically created during the evaluation run of model DreadPoor/Blunt_Edge-8B-SLERP The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Blunt_Edge-8B-SLERP-details.tabular10K<n<100K0 likes12 downloads2y agoHugging Face26Edgerunners /Thalia-Functions-13kBetter approach to Edgerunners/glaive-first-human-only-deduped, containing general questions a human might ask an AI assistant. Best performers I went through all OpenRouter offered models with theme-based filtering and then landed on a few top performers: koboldai/psyfighter-13b-2 (72) intel/neural-chat-7b (54) mistralai/mistral-small (42) openchat/openchat-7b (37) gryphe/mythomax-l2-13b:nitro (31) mistralai/mistral-tiny (31) qwen/qwen-7b-chat (29) the number representing how often it… See the full description on the dataset page: https://huggingface.co/datasets/Edgerunners/Thalia-Functions-13k.text10K<n<100K0 likes11 downloads2y agoHugging Face27Edgerunners /Thalia-Functions-OC3.6-33kSame idea as Edgerunners/Thalia-Functions-13k except instead of running through all OpenRouter models and then also generating locally, I have used the newly released OpenChat 3.6 which performed the best on its own, semantically deduped. Topics covered: General Trivia and Entertainment Questions Lifestyle, Self-Improvement, and Productivity Questions Music and Media Requests Financial and Economic Questions Technology and Computer-Related Questions Food, Cooking, and Grocery Shopping… See the full description on the dataset page: https://huggingface.co/datasets/Edgerunners/Thalia-Functions-OC3.6-33k.text10K<n<100K0 likes11 downloads2y agoHugging Face28neoneye /simon-arc-solve-edge-v3 Version 1 ARC-AGI Tasks where the job is to identify where the top/bottom/left/right edge of the object is located. example count: 4-5. test count: 1-2. image size: 3-5. Version 2 image size: 3-10. Version 3 image size: 3-5. Focus on identifying diagonal edges. textimage-to-text100K<n<1M0 likes11 downloads2y agoHugging Face29neoneye /simon-arc-solve-edge-v4 Version 1 ARC-AGI Tasks where the job is to identify where the top/bottom/left/right edge of the object is located. example count: 4-5. test count: 1-2. image size: 3-5. Version 2 image size: 3-10. Version 3 image size: 3-5. Focus on identifying diagonal edges. Version 4 image size: 3-10. Focus on identifying diagonal edges. textimage-to-text100K<n<1M0 likes10 downloads2y agoHugging Face30neoneye /simon-arc-solve-edge-v5 Version 1 ARC-AGI Tasks where the job is to identify where the top/bottom/left/right edge of the object is located. example count: 4-5. test count: 1-2. image size: 3-5. Version 2 image size: 3-10. Version 3 image size: 3-5. Focus on identifying diagonal edges. Version 4 image size: 3-10. Focus on identifying diagonal edges. Version 5 image size: 3-10. Enabled all edge_names: top_left, top, top_right, left, right, bottom_left, bottom… See the full description on the dataset page: https://huggingface.co/datasets/neoneye/simon-arc-solve-edge-v5.textimage-to-text100K<n<1M0 likes10 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.