Horizontal proactivity
Pursuing a need the current state already names. A customer gives a name and a ZIP code, so the agent looks up the account.
Preprint · 2026
1IBM2Weizmann Institute of Science
Tool-using LLM agents do what they are asked, yet a task often needs information the user never mentions. We study what an agent should pursue unasked, measure it against need graphs with no model judge, and train it with Q&D, which learns from the consequences of its own questions.
of the required evidence recovered on MuSiQue at equal retrieval spend: the same 8B model, prompted and then trained
multi-hop benchmarks where, at equal retrieval spend, the 8B questioner outperforms GPT-OSS-120B, a prompted model 15× larger in the same role
task success in τ²-bench retail, with no further training: it asks less and finds more
An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what information it should pursue. We study a distinct axis of proactivity: its content. Horizontal proactivity pursues unstated information that the current context already identifies, and vertical proactivity pursues needs that only earlier evidence reveals. A need graph, recovered from a benchmark's own decomposition, records which needs depend on which, so both forms, and whether the agent stops at the right time, can be scored from a transcript without a model judge. To learn this behavior, we propose Q&D (questioner and drafter), which trains a questioner to prefer the question whose continuation retrieves more of the required evidence, with no reward model or judge. On held-out splits of three multi-hop question-answering benchmarks, at equal retrieval spend, the trained questioner improves both forms of proactivity over the same model, prompted, and outperforms a prompted model 15× larger in the same role on two of the three, and the gain persists after controlling for question volume and length. Without further training, we place the questioner in an interactive customer-service agent with a simulated customer, where it completes more tasks while asking fewer questions, and in retail it outperforms the 15× larger model with fewer follow-up turns from the customer. These results show that proactivity depends not only on whether an agent acts without being asked, but also on what it chooses to pursue and when it stops.
Work on proactive agents mostly asks whether and when an agent should act on its own. We ask what it should go after, and distinguish two forms.
Pursuing a need the current state already names. A customer gives a name and a ZIP code, so the agent looks up the account.
Pursuing a need that only newly found evidence names. The account lists the order, the order names the product, and the product lists the size-8 variant.
A need graph records the evidence a task requires, with an edge wherever one need can be named only after another is found. Multi-hop question-answering benchmarks supply these graphs through their own decompositions of their questions. Both forms of proactivity, and whether the agent stops at the right time, are then scored from a transcript, with no model judge.
Agents are compared at equal retrieval spend, after the same number of questions, so asking more cannot pass for asking better.
Q&D (questioner and drafter) splits the agent in two. A questioner asks one question at a time or stops, a retriever answers it from the task's evidence, and a frozen drafter folds the evidence into a draft. Because the drafter is fixed, every change in what the agent holds is caused by a question.
To learn, Q&D forks a run at one state, continues it after several candidate questions and after stopping, and prefers the candidate whose continuation retrieves more of the required evidence. Asking is preferred over stopping while required evidence is still missing. There is no reward model and no model judge. The questioner is trained in three stages:
On held-out test splits: MuSiQue at equal retrieval spend, and τ²-bench under the benchmark's own per-dialogue caps.
| Setting | Measure | Qwen3-8B, prompted | Q&D questioner |
|---|---|---|---|
| MuSiQue | Required evidence recovered | 78% | 90% |
| τ²-bench retail, base prompt | Task success | 13% | 34% |
| τ²-bench retail, stop prompt | Task success | 12% | 32% |
The questioner is a LoRA adapter on Qwen3-8B, released with both training seeds. It reads the prompt template it was trained on and replies with one JSON action per step, to ask or to stop. The model card has a complete example, run on a GPU with its output shown.
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B")
base = AutoModelForCausalLM.from_pretrained(
"Qwen/Qwen3-8B", dtype=torch.bfloat16, device_map="auto"
)
questioner = PeftModel.from_pretrained(base, "dolev31/ProactiveInquirer-Qwen3-8B")
To run it on your own machine, with GGUF files for Ollama, LM Studio and llama.cpp, or with merged weights that need no PEFT:
ollama run hf.co/dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M --think=false
To run the whole loop with a retriever, a drafter and the paper's metrics, start from the repository: make smoke runs it end to end on a synthetic suite, with no API keys.
A tool-using LLM agent usually does what the user asks, yet completing the task often needs information the user never mentions. Work on proactive agents mostly studies whether and when an agent should act on its own. This paper studies what the agent pursues unasked, in a horizontal and a vertical form, and measures and trains both.
Horizontal proactivity pursues a need the current state already names, such as looking up an account from a name and ZIP code the customer gave. Vertical proactivity pursues a need that only newly found evidence names, such as the order that account lists. Vertical needs form chains, and each link can be pursued only after the one before it is found.
Each task carries a need graph recovered from the benchmark's own decomposition of its question. A finished transcript is scored by which required evidence it retrieved, matched by identifier, at which depth, and whether the agent stopped at the right time. No quantity the paper reports is judged by a model.
From the consequences of its questions. A recorded run is forked at a step, eight alternative questions are sampled there, and each is continued to the end. Pairs are ordered by whether the run answered the task, then by which reached the complete evidence sooner, then by how much evidence each turn added. The questioner is trained in three stages: imitation of good decisions, direct preference optimization on question pairs, and direct preference optimization on question pairs together with contrasts that rank asking above stopping at unfinished states.
Search agents trained with outcome rewards credit a whole run with whether its final answer was right. Q&D credits each question with what followed it, so a question that reaches an unstated need early wins even when the final answers agree. Answers decide only 6% of the pairs Q&D trains on.
Work on clarifying questions learns what to ask the user when a request is ambiguous or incomplete. Here the questioner goes after the missing information itself, in the evidence it can retrieve. In a customer-service agent it completes more retail tasks than GPT-OSS-120B, a prompted model 15× larger, with fewer follow-up turns from the customer, which the paper attributes to the agent finding more of what the task needs in the store's records.
Training data are mined from MuSiQue, StrategyQA and 2WikiMultiHopQA, and results are read on held-out test splits of 200 tasks per benchmark. The trained questioner is Qwen3-8B. GPT-OSS-120B is the drafter and the answerer in every arm except the single-model control, and a prompted comparator in the questioner's role. The transfer study uses the retail and airline domains of τ²-bench with a simulated customer.
At equal retrieval spend, the trained 8B questioner recovers more of the required evidence than the same model, prompted, on every suite, and reaches deeper into each task's dependencies. On MuSiQue it recovers 90% of the required evidence against 78%. It leads GPT-OSS-120B, a prompted model 15× larger in the same role, on MuSiQue and StrategyQA, by 7.0 and 4.3 points, and trails it on 2WikiMultiHopQA, by 3.5 points. The gain comes from what it asks, not from asking more or longer questions. Stopping on its own, it asks fewer questions than either prompted model and still leads on the deepest chains.
Yes, without further training. Placed in a customer-service agent on τ²-bench, it raises retail task success from 13% and 12% to 34% and 32% under two prompt variants, better on ten of twenty-five tasks and worse on none in each. It asks less and finds more: 1.3 and 1.4 fewer questions per retail dialogue, more of which reach the records the task needs. In airline, task success rises by 6.9 and 6.4 points. Against GPT-OSS-120B in retail, it completes 17.0 and 12.7 points more tasks, with 1.6 and 1.8 fewer follow-up turns from the customer.
Not yet detectably. The extra evidence does not yet reach the final answers, and training teaches what to ask more readily than when to stop.
The two reported seeds share one supervised checkpoint and one pair export, so intervals are over tasks alone. The released retrieval pools are small enough that most prerequisite edges do not gate retrieval, so the vertical result concerns depth in a decomposition. Every user-facing number comes from a simulated customer, and Q&D is one round of off-policy training, not compared with on-policy reinforcement learning.
Proactive and mixed-initiative agents, benchmarks that test whether a model recognizes missing information or unstated needs, learning what to ask, including clarifying questions, agentic search and retrieval-augmented generation, question decomposition for multi-hop question answering, step-level preference training, and simulated users for evaluating agents.
The code, with the agent loop, the need-graph builders and metrics, the training ladder, the analyses and a test suite, is on GitHub at dolev31/ProactiveInquirer. The trained questioner is on Hugging Face at dolev31/ProactiveInquirer-Qwen3-8B, a LoRA adapter with both training seeds, with merged weights and GGUF quantizations for Ollama and LM Studio in companion repositories. All are under Apache-2.0.
@article{levy2026asking,
title = {Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents},
author = {Levy, Ido and Yehudai, Asaf and Shlomov, Segev and Adi, Asaf and Choshen, Leshem},
journal = {arXiv preprint},
year = {2026}
}