datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
recall-rewrite-oasst1
Recall Rewrite OASST1: knowledge-aligned SFT data
Data release for the paper "Stick to What You Know: A Study of Knowledge-Aligned Supervised Fine-Tuning"
(Becker, Kemmler, Thulke, Schäfer, Dugast, Ney; accepted at EMNLP 2026, Main Conference).
Knowledge-aligned SFT constrains supervised fine-tuning targets to what the base model already knows.
Recall Rewrite implements this without external evidence: every gold response of the SFT set is
decomposed into atomic claims, each… See the full description on the dataset page: https://huggingface.co/datasets/apptek-com/recall-rewrite-oasst1.recall-radar
Cross-Border Product Recall Dataset
A structured, machine-readable dataset of consumer-facing product recalls in three high-stakes categories (cosmetic, baby_product, food) across three regions (US, EU, JP), drawn from four authoritative agencies. Designed for cross-border e-commerce sellers and AI shopping assistants.
Overview
Product recalls are scattered across CPSC's web table, FDA's RSS feed, Safety Gate's search interface, and 消費者庁's Japanese-language… See the full description on the dataset page: https://huggingface.co/datasets/Aulvem/recall-radar.agentic-state-recall-v1
Agentic State Recall v1 (ENERZAi 내부, 2026-09-22)
실제 에이전트 궤적(AgentTuning alfworld·webshop, nebius SWE-agent; neulab/agent-data-collection 표준화본)에서 개체별 상태 변화를 프로그램으로 복원해 구조화 정답을 만들고,
Qwen3.8-27B 가 자연어로 문장화한 뒤 역파싱 검증·근거 게이트·번호 정합 게이트를 통과한 문항만 남긴 학습 데이터. AMA-Bench 의 네 유형(A 회상 · B 인과 · C 상태 갱신 · D 상태 추상화)을 겨냥한다.
B 유형의 "왜"·"실패 뒤 다음 행동" 문항은 궤적에 기록된 에이전트의 이유(Thought) 를 근거로 한다.
프롬프트는 AMA compaction_v3_nostate 하네스 형식(Task / Step Index / Most Recent(Thought: 줄 포함) / Recalled / Questions /… See the full description on the dataset page: https://huggingface.co/datasets/HBKenerzai/agentic-state-recall-v1.
