decomposition
multihop-question-decompositiondaily-paper-2026-07-09-retriever-vs-decomposition-skill-routing
Retriever Bottleneck vs. Decomposition
TL;DR — In a real 1,898-skill bilingual agent harness, the ceiling-gap diagnostic
shows that neither decomposition nor a better retriever raises routing coverage: the
binding constraint is skill-corpus redundancy. Decomposition still earns its place on
execution ordering.
ThakiCloud AI Research · 2026-07-09 (v2, revised) · 📝 Tech blog (KO)
Problem
Operators of large agent harnesses that route natural-language requests to… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-09-retriever-vs-decomposition-skill-routing.task168_strategyqa_question_decomposition
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task168_strategyqa_question_decomposition
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task168_strategyqa_question_decomposition.satsec-decomposition
SatSec Grounded Objective-Decomposition Dataset
Version 2.0 is a leakage-controlled replacement for the original dataset used in
A Controlled Candidate-Set Benchmark for Offline Satellite-Security Plan
Decomposition
(DOI 10.48550/arXiv.2607.26371).
It contains 24 authored full decompositions and 83 mechanically derived next-step rows
across 24 cases.
There are 82 train rows and 25 test rows; the six test cases never occur in training.
Important v2 correction
The… See the full description on the dataset page: https://huggingface.co/datasets/paolocmo/satsec-decomposition.question-decomposition-peft
Question Decomposition Dataset for PEFT/LoRA Training
This dataset is designed to fine-tune language models to decompose complex multi-hop questions into simpler, sequential subquestions — NOT to answer them. The goal is to teach models the reasoning structure needed to break down complex queries following a canonical syntactic decomposition before retrieval or answering.
This dataset then aims at enhancing the syntactic understanding of questions to further decompose them. Most… See the full description on the dataset page: https://huggingface.co/datasets/Anvix/question-decomposition-peft.DeCompBench
DecompBench
DecompBench is a benchmark for evaluating decomposition attacks on tool-using LLM agents. A harmful task is split into subtasks that are individually benign, and each subtask is issued to the agent without the conversation history of the others. Tasks execute against real services (GitLab, OwnCloud, RocketChat, PostgreSQL, Redis, Notion, Plane) in a Docker environment, and success is measured by checkpoints that inspect the resulting environment state.… See the full description on the dataset page: https://huggingface.co/datasets/decompositionbench/DeCompBench.
