Anthropic's book-swap test: agents' weak spot was reading tastes
In Project Swap, 201 Anthropic employees let Claude agents trade physical books for them. Most of the shortfall came from misjudged preferences rather than bargaining, and stronger models won.

Key takeaways
- Claude agents traded real books for 201 Anthropic staff, who got about their 5th choice of 10.
- Misread preferences caused 85% of the shortfall; bargaining on the floor caused the other 15%.
- On mixed trading floors, agents running Opus always came out ahead of those on Haiku.
Anthropic has published the results of Project Swap, an experiment in which 201 of its employees let Claude agents trade real, physical books on their behalf. The study, posted on Thursday, September 24, found that the agents were better at bargaining than at knowing what their owners wanted.
Each participant brought a book to give away and described their reading tastes in a short chat with Claude, and the agents then met on trading floors in six office pools. On average, people ended up at 0.55 on their own rankings, roughly their fifth choice from a list of ten. The best possible assignment would have scored 0.89, roughly their second choice.
Most of the loss came before any bargaining
Anthropic split the shortfall in two. Working from Claude's imprecise rankings accounted for 85% of the gap, and the free-for-all trading floor for the remaining 15%. From a five-minute chat, an agent's ranking of books matched its person's on 61% of pairs. Guessing would get you 50%. Collaborative filtering, the "people who liked X also liked Y" method behind many recommendation engines, reached about 55%. Longer conversations helped: doubling the words in the intake was associated with 4.1 percentage points more agreement.
That is the finding most discussion of shopping agents skips. The debate tends to be about whether agents will be tough negotiators or easy marks. In this test the bigger problem sat upstream: the agent was bargaining hard for a picture of its owner that was only partly right.
Stronger models won, and the losers may not notice
Model strength still mattered. On Claude's rankings, agents on all-Haiku floors averaged 0.75 and agents on all-Opus floors averaged 0.88. On floors split half Opus and half Haiku, "the Opus agents always came out ahead".
That repeats what Anthropic saw in Project Deal, its earlier experiment, in which 69 employees let agents buy and sell personal items in December 2025. There, an item sold by an Opus agent went for $3.64 more on average than one sold by a Haiku agent, yet people rated the fairness of their deals almost identically, 4.05 for Opus deals and 4.06 for Haiku deals on a seven-point scale. If the people with the weaker agent cannot feel the difference, they have no reason to pay for a better one.
The instructions given to an agent mattered less than its model. An agent told to be "ruthless" scored about 0.02 higher than a prosocial one on the same floor, and prosocial agents accepted a book lower on their own ranking twice as often, though both cases were rare. The agents were also mostly honest: between 78% and 96% mentioned their top pick at some point, and only about 1 in 100 of those lied about it.
What the humans made of it
Among those who answered the survey, average satisfaction with the book received was 7.2 out of 10. Asked what share of their annual book budget they would hand to an agent, the average answer was about 30%. One participant was unhappy: "I didn't like being on the docile side of the experiment". Another read the full log of their agent's trades, found a book it had held and then swapped away, and bought it. Their verdict: "For agent economies, observability about the process will be as important as the outcome, [as] this gives people recourse".
The limits Anthropic states
This was Anthropic staff, trading books, through agents built on Anthropic's own models. The post says so: its employees "are probably more eager to trust Claude than most people", and all agents were "post-trained to be polite and largely cooperative". Some participants never brought their books in, and there was no system to track pick-ups.
The open question is the one the experiment could not test. When the agent across the table is built by someone else, and built to exploit yours, does the 15% lost in bargaining stay at 15%?
- Anthropic
- AI research
- AI agents
- Claude
- Agent commerce
Sources
- Project Swap: What happens when agents trade for us? — Anthropic, Sep 24, 2026
- Project Deal: our Claude-run marketplace experiment — Anthropic, Apr 24, 2026
Related stories

Anthropic says Claude agents found a new enzyme system in 21 hours
About 950 Claude agents combed a DNA database and picked out a reverse transcriptase sitting beside a CRISPR-like array of repeats. Anthropic's human scientists have started characterising it in the lab. What it does is still unknown.
4 min read

Mathematicians solve the last sporadic Galois case with AI help
A six-author paper shows the Mathieu group M23 is a Galois group over the rationals, closing a list open since the 1980s. The acknowledgments credit Claude and ChatGPT; the text, the authors say, is all human.
3 min read

First industry count finds about 7,000 humanoid robots sold in 2025
The International Federation of Robotics says many of them were bought to collect data for AI models, not to do productive work.
3 min read
Comments
No comments yet. Start the conversation.