datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/BlueIsGreen/Edge-Agent-Reasoning-WebSearch-260K.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/DEMIRUNC/Edge-Agent-Reasoning-WebSearch-260K.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/Torenn/Edge-Agent-Reasoning-WebSearch-260K.edge-agent-reasoning-websearch-260k
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/ppenner/edge-agent-reasoning-websearch-260k.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/JACKYS999/Edge-Agent-Reasoning-WebSearch-260K.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/svryn/Edge-Agent-Reasoning-WebSearch-260K.EdgeReason
EdgeReason
EdgeReason is a compact verifier-backed dataset for improving small and
edge-deployable language models on tool use, structured JSON outputs,
state/table/unit/date reasoning, compact Mathlib-derived SFT, and routing
between direct answer, tool use, retrieval, clarification, and escalation.
The dataset is designed for teams training small models with SFT, DPO, RLVR,
GRPO, rejection sampling, and internal evaluation loops. It is not tied to any
model vendor or… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/EdgeReason.memaudit-data
MemAudit Dataset Artifact
This dataset contains finite memory-writing packages, exact/verified solver outputs, model-adjudicated natural support-slice packages, human-edited fictional seed packages, and exported-system diagnostic results for Mem0, Letta, and A-Mem. It is intended to support the NeurIPS 2026 E&D submission MemAudit: An Exact-Oracle Evaluation Protocol for Budgeted Long-Term LLM Memory Writing.
Historical directory names may contain oraclemem; paper-facing terminology… See the full description on the dataset page: https://huggingface.co/datasets/edgeclustr/memaudit-data.EDGE-DatasetThis is the dataset repository of paper EDGE: Enhanced Grounded GUI Understanding with Enriched Multi-Granularity Synthetic Data.
Considering the huge number of images, the all_items.jsonl provided here contains the final QA pairs for training, but does not contain images for the time being. We will release all images as soon as possible.
You can also follow the mark_webpages and dataset.py scripts provided in the code repository to generate your own webpage image-question-answering dataset.… See the full description on the dataset page: https://huggingface.co/datasets/EDGEwww25/EDGE-Dataset.CRaQAn_v1
Coreference Resolution in Question Answering (CRaQAn)
250+ question-answer pairs that require coreference resolution across sentences from selected Wikipedia passages.
Generation Process
Given the relative complexity of our task (coreference resolution across passages for question-answering), we aimed
to avoid crowd-sourcing this dataset and instead focused on using LLMs to automate our process. Every question-answer
pair in the CRaQAn dataset was automatically generated… See the full description on the dataset page: https://huggingface.co/datasets/Edge-Pyxos/CRaQAn_v1.wind-edge-1.6-sft
Wind Lite SFT
Custom supervised fine-tuning dataset for Wind Lite 1.6 by North AI.
Dataset Summary
20,000 high-quality instruction-response pairs covering identity grounding, math reasoning, coding, general knowledge, and multi-turn conversations.
Data Composition
Category
Count
Description
Math & Reasoning
~7,000
Arithmetic, algebra, percentages, unit conversions — with step-by-step working
Coding
~4,000
Python, JavaScript, SQL, systems — with… See the full description on the dataset page: https://huggingface.co/datasets/North-ML1/wind-edge-1.6-sft.edgeqa-resource-release
EdgeQA resource release (OLP + OSP)
This Hugging Face dataset repository contains a release-ready snapshot of the EdgeQA resource family:
EdgeQA (grounded QA exports at multiple budgets), EdgeCoverBench (robustness + abstention stress tests), and BEIR-style IR test collections.
Links
Code: https://github.com/cactusYuri/edgeqa
Author: Yuli Zhang (Beijing University of Posts and Telecommunications) — zhangyuli@bupt.edu.cn
Contents
corpora/: redistributable… See the full description on the dataset page: https://huggingface.co/datasets/cactusYuri/edgeqa-resource-release.
