• Home
  • News
  • Blog
  • Releases
  • LLM history
  • Compare LLMs
  • Library
  • About
⌘K
Sign in

A blog and notes on development. The easiest way to reach me is via the social links below.

Contacts
talalaev.misha@gmail.com
Documents
Personal data processing policyPersonal data processing consent
Photo: National Cancer Institute / Unsplash

An LLM agent closes the deal — and can still lose money

Sh0ny
Sh0ny
11 августа 2026
  1. Home
  2. Blog
  3. An LLM agent closes the deal — and can still lose money
2 min read

In short

A study of 9,840 negotiations shows that an agent's ability to reach agreement does not mean the contract is profitable or reliable. The outcome is driven most of all by the supplier's model and by how strategically patient the prompt tells it to be.

An autonomous agent can almost always bring a negotiation to an agreement, but that does not guarantee a good deal. In the experiment agents reached agreement in 98.9% of cases and captured 95.4% of the theoretically available gains — though surplus rounds of negotiation ate another 21–34% of that result.

That is an important shift in the criteria of evaluation. For a procurement agent it is not enough to check whether it can bargain and find a compromise. You have to look at how much time it spends reaching agreement, whether it accepts loss-making terms, and who ends up with the value created.

In the study nine LLMs from the OpenAI, Google and Alibaba ecosystems conducted 9,840 negotiations with each other. The average agent needed 2.98 rounds against 1.25 for the benchmark equilibrium model. That is, an outwardly successful dialogue can be markedly worse than short, rational bargaining.

Moreover, the distribution of gains proved tied not so much to capability rankings as to the model's provenance. In self-play the buyer on average took 40% with OpenAI, 50% with Google and 70% with Alibaba's Qwen. Change the supplier and the share shifted by 7–18 percentage points. The choice of vendor here becomes not only a technical but an economic decision.

The prompt too can change the result more than the model's "cleverness". The authors distinguish the business owner's economic patience from the strategic patience set for the agent in its instructions: it was that free parameter that explained 90% of the variation in how gains were split.

The study's limitations are substantial as well. This concerns a canonical supply chain problem: the buyer knows their demand, the seller does not, and the sides bargain over quantity and payment. It is not a full test of real procurement. Besides, base models accepted individually unprofitable contracts in 19.2% of cases; for mid-tier and flagship models the figure was 0–0.6%. So an automatic profitability check is a mandatory safeguard, especially for weaker models.

The practical conclusion is simple: before deploying an agent you have to test not only average gains but latency, worst-case contracts and any skew in one side's favour. And record the prompt separately: changing "patience" can redistribute money without a single model update.

If you delegated procurement to an agent, which would you consider more dangerous: a few surplus rounds of bargaining, or a systematic skew of the deal in the supplier's favour? Source: cs.AI updates on arXiv.org

новостиaiагентыllm
More AI-tool write-ups on the Telegram channel — short and to the point
Subscribe on Telegram

Comments

(0)
​