Preprint · 2026

Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents

Ido Levy1,2· Asaf Yehudai1· Segev Shlomov1· Asaf Adi1· Leshem Choshen1,2

1IBM2Weizmann Institute of Science

Figure 1. One customer request handled by the same model, prompted and trained with Q&D. The customer asks to change a pending order's boots to size 8 and gives a name and ZIP code, not the email or the order ID. Prompted, the agent spends all 16 of its questions asking the customer for the email or the order ID and the task fails. Trained, it finds the user profile from the name and ZIP code (horizontal), then the order, the boots and the size-8 variant from values it found (vertical), and the task is completed. Panel b shows Q&D: a questioner asks or stops, a retriever answers, a frozen drafter updates the draft, and training prefers the question whose continuation finds more of what the task needs.
One request, two agents. A customer asks to change the boots in a pending order to size 8 and gives a name and a ZIP code, but not the email or the order number. The same model, prompted, asks the customer for them sixteen times and the task fails. Trained with Q&D, it finds the account, the order and the size-8 boots in the store's records and completes the task: a horizontal step from what the customer said, then vertical steps from what it found.

Tool-using LLM agents do what they are asked, yet a task often needs information the user never mentions. We study what an agent should pursue unasked, measure it against need graphs with no model judge, and train it with Q&D, which learns from the consequences of its own questions.

78%→90%

of the required evidence recovered on MuSiQue at equal retrieval spend: the same 8B model, prompted and then trained

2 of 3

multi-hop benchmarks where, at equal retrieval spend, the 8B questioner outperforms GPT-OSS-120B, a prompted model 15× larger in the same role

13%→34%

task success in τ²-bench retail, with no further training: it asks less and finds more

Abstract

An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what information it should pursue. We study a distinct axis of proactivity: its content. Horizontal proactivity pursues unstated information that the current context already identifies, and vertical proactivity pursues needs that only earlier evidence reveals. A need graph, recovered from a benchmark's own decomposition, records which needs depend on which, so both forms, and whether the agent stops at the right time, can be scored from a transcript without a model judge. To learn this behavior, we propose Q&D (questioner and drafter), which trains a questioner to prefer the question whose continuation retrieves more of the required evidence, with no reward model or judge. On held-out splits of three multi-hop question-answering benchmarks, at equal retrieval spend, the trained questioner improves both forms of proactivity over the same model, prompted, and outperforms a prompted model 15× larger in the same role on two of the three, and the gain persists after controlling for question volume and length. Without further training, we place the questioner in an interactive customer-service agent with a simulated customer, where it completes more tasks while asking fewer questions, and in retail it outperforms the 15× larger model with fewer follow-up turns from the customer. These results show that proactivity depends not only on whether an agent acts without being asked, but also on what it chooses to pursue and when it stops.

Two forms of proactivity

Work on proactive agents mostly asks whether and when an agent should act on its own. We ask what it should go after, and distinguish two forms.

Horizontal proactivity

Pursuing a need the current state already names. A customer gives a name and a ZIP code, so the agent looks up the account.

name + ZIP code→account

Vertical proactivity

Pursuing a need that only newly found evidence names. The account lists the order, the order names the product, and the product lists the size-8 variant.

account→order→product→size-8 variant

Measured without a model judge

A need graph records the evidence a task requires, with an edge wherever one need can be named only after another is found. Multi-hop question-answering benchmarks supply these graphs through their own decompositions of their questions. Both forms of proactivity, and whether the agent stops at the right time, are then scored from a transcript, with no model judge.

Agents are compared at equal retrieval spend, after the same number of questions, so asking more cannot pass for asking better.

Q&D: learning from the consequences of its own questions

Q&D (questioner and drafter) splits the agent in two. A questioner asks one question at a time or stops, a retriever answers it from the task's evidence, and a frozen drafter folds the evidence into a draft. Because the drafter is fixed, every change in what the agent holds is caused by a question.

To learn, Q&D forks a run at one state, continues it after several candidate questions and after stopping, and prefers the candidate whose continuation retrieves more of the required evidence. Asking is preferred over stopping while required evidence is still missing. There is no reward model and no model judge. The questioner is trained in three stages:

  1. Imitation of good decisions: stop where the required evidence is already in hand, and otherwise ask the best sampled question.
  2. Direct preference optimization on pairs of questions asked from the same state.
  3. The same, with contrasts between asking and stopping added.

Results

On held-out test splits: MuSiQue at equal retrieval spend, and τ²-bench under the benchmark's own per-dialogue caps.

SettingMeasureQwen3-8B, promptedQ&D questioner
MuSiQueRequired evidence recovered78%90%
τ²-bench retail, base promptTask success13%34%
τ²-bench retail, stop promptTask success12%32%
  • It improves both forms of proactivity over the same model, prompted, on held-out splits of MuSiQue, StrategyQA and 2WikiMultiHopQA.
  • The gain comes from what it asks, not from asking more or longer questions: it holds against a question-volume control and a length control.
  • In τ²-bench retail it is better on ten of twenty-five tasks and worse on none, under each of two prompts. It asks less and finds more: 1.3 and 1.4 fewer questions per dialogue, more of which reach the records the task needs. In airline, success rises by 6.9 and 6.4 points.
  • Against GPT-OSS-120B in retail, it completes 17.0 and 12.7 points more tasks, with 1.6 and 1.8 fewer follow-up turns from the customer.
Required-evidence coverage against mean retrieval calls on MuSiQue and StrategyQA, and answer correctness on FRAMES. The trained questioner, labeled by its budget of 4 to 24 calls, is compared with the same model and GPT-OSS-120B, both prompted.
Required-evidence coverage against the mean number of retrieval calls (MuSiQue, StrategyQA) and answer correctness (FRAMES): the trained questioner, labeled by its call budget, against the same model and GPT-OSS-120B, both prompted.
Differences on tau2-bench, the trained questioner minus each prompted baseline, in retail and airline under two prompts: task success in percentage points on the left, follow-up turns per dialogue on the right, where negative means fewer turns.
τ²-bench, the trained questioner minus each prompted baseline: task success in percentage points (left) and follow-up turns from the customer per dialogue (right, fewer to the left).

Use the trained questioner

The questioner is a LoRA adapter on Qwen3-8B, released with both training seeds. It reads the prompt template it was trained on and replies with one JSON action per step, to ask or to stop. The model card has a complete example, run on a GPU with its output shown.

import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B")
base = AutoModelForCausalLM.from_pretrained(
    "Qwen/Qwen3-8B", dtype=torch.bfloat16, device_map="auto"
)
questioner = PeftModel.from_pretrained(base, "dolev31/ProactiveInquirer-Qwen3-8B")

To run it on your own machine, with GGUF files for Ollama, LM Studio and llama.cpp, or with merged weights that need no PEFT:

ollama run hf.co/dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M --think=false

To run the whole loop with a retriever, a drafter and the paper's metrics, start from the repository: make smoke runs it end to end on a synthetic suite, with no API keys.

Questions

What does it mean for an LLM agent to be proactive about content?

A tool-using LLM agent usually does what the user asks, yet completing the task often needs information the user never mentions. Work on proactive agents mostly studies whether and when an agent should act on its own. This paper studies what the agent pursues unasked, in a horizontal and a vertical form, and measures and trains both.

What is the difference between horizontal and vertical proactivity?

Horizontal proactivity pursues a need the current state already names, such as looking up an account from a name and ZIP code the customer gave. Vertical proactivity pursues a need that only newly found evidence names, such as the order that account lists. Vertical needs form chains, and each link can be pursued only after the one before it is found.

How is proactivity measured without an LLM judge?

Each task carries a need graph recovered from the benchmark's own decomposition of its question. A finished transcript is scored by which required evidence it retrieved, matched by identifier, at which depth, and whether the agent stopped at the right time. No quantity the paper reports is judged by a model.

How does Q&D train a questioner without a reward model?

From the consequences of its questions. A recorded run is forked at a step, eight alternative questions are sampled there, and each is continued to the end. Pairs are ordered by whether the run answered the task, then by which reached the complete evidence sooner, then by how much evidence each turn added. The questioner is trained in three stages: imitation of good decisions, direct preference optimization on question pairs, and direct preference optimization on question pairs together with contrasts that rank asking above stopping at unfinished states.

How does this differ from training agentic search on answer correctness?

Search agents trained with outcome rewards credit a whole run with whether its final answer was right. Q&D credits each question with what followed it, so a question that reaches an unstated need early wins even when the final answers agree. Answers decide only 6% of the pairs Q&D trains on.

How does this differ from asking clarifying questions?

Work on clarifying questions learns what to ask the user when a request is ambiguous or incomplete. Here the questioner goes after the missing information itself, in the evidence it can retrieve. In a customer-service agent it completes more retail tasks than GPT-OSS-120B, a prompted model 15× larger, with fewer follow-up turns from the customer, which the paper attributes to the agent finding more of what the task needs in the store's records.

Which benchmarks and models does the paper use?

Training data are mined from MuSiQue, StrategyQA and 2WikiMultiHopQA, and results are read on held-out test splits of 200 tasks per benchmark. The trained questioner is Qwen3-8B. GPT-OSS-120B is the drafter and the answerer in every arm except the single-model control, and a prompted comparator in the questioner's role. The transfer study uses the retail and airline domains of τ²-bench with a simulated customer.

What are the main results?

At equal retrieval spend, the trained 8B questioner recovers more of the required evidence than the same model, prompted, on every suite, and reaches deeper into each task's dependencies. On MuSiQue it recovers 90% of the required evidence against 78%. It leads GPT-OSS-120B, a prompted model 15× larger in the same role, on MuSiQue and StrategyQA, by 7.0 and 4.3 points, and trails it on 2WikiMultiHopQA, by 3.5 points. The gain comes from what it asks, not from asking more or longer questions. Stopping on its own, it asks fewer questions than either prompted model and still leads on the deepest chains.

Does the questioner work beyond question answering?

Yes, without further training. Placed in a customer-service agent on τ²-bench, it raises retail task success from 13% and 12% to 34% and 32% under two prompt variants, better on ten of twenty-five tasks and worse on none in each. It asks less and finds more: 1.3 and 1.4 fewer questions per retail dialogue, more of which reach the records the task needs. In airline, task success rises by 6.9 and 6.4 points. Against GPT-OSS-120B in retail, it completes 17.0 and 12.7 points more tasks, with 1.6 and 1.8 fewer follow-up turns from the customer.

Does finding more evidence give better final answers?

Not yet detectably. The extra evidence does not yet reach the final answers, and training teaches what to ask more readily than when to stop.

What are the limitations?

The two reported seeds share one supervised checkpoint and one pair export, so intervals are over tasks alone. The released retrieval pools are small enough that most prerequisite edges do not gate retrieval, so the vertical result concerns depth in a decomposition. Every user-facing number comes from a simulated customer, and Q&D is one round of off-policy training, not compared with on-policy reinforcement learning.

Which research areas does this work connect to?

Proactive and mixed-initiative agents, benchmarks that test whether a model recognizes missing information or unstated needs, learning what to ask, including clarifying questions, agentic search and retrieval-augmented generation, question decomposition for multi-hop question answering, step-level preference training, and simulated users for evaluating agents.

What is released, and where?

The code, with the agent loop, the need-graph builders and metrics, the training ladder, the analyses and a test suite, is on GitHub at dolev31/ProactiveInquirer. The trained questioner is on Hugging Face at dolev31/ProactiveInquirer-Qwen3-8B, a LoRA adapter with both training seeds, with merged weights and GGUF quantizations for Ollama and LM Studio in companion repositories. All are under Apache-2.0.

Citation

@article{levy2026asking,
  title   = {Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents},
  author  = {Levy, Ido and Yehudai, Asaf and Shlomov, Segev and Adi, Asaf and Choshen, Leshem},
  journal = {arXiv preprint},
  year    = {2026}
}