Skip to content
nanoai

Anthropic's book-swap test: agents' weak spot was reading tastes

In Project Swap, 201 Anthropic employees let Claude agents trade physical books for them. Most of the shortfall came from misjudged preferences rather than bargaining, and stronger models won.

By Siva Charakani3 min read

AI research graphic showing a Claude-like AI robot reading books in a bookstore, with an experiment dashboard comparing AI recommendations and human reading preferences from 201 staff participants.
Image: AI Generated

Key takeaways

  • Claude agents traded real books for 201 Anthropic staff, who got about their 5th choice of 10.
  • Misread preferences caused 85% of the shortfall; bargaining on the floor caused the other 15%.
  • On mixed trading floors, agents running Opus always came out ahead of those on Haiku.

Anthropic has published the results of Project Swap, an experiment in which 201 of its employees let Claude agents trade real, physical books on their behalf. The study, posted on Thursday, September 24, found that the agents were better at bargaining than at knowing what their owners wanted.

Each participant brought a book to give away and described their reading tastes in a short chat with Claude, and the agents then met on trading floors in six office pools. On average, people ended up at 0.55 on their own rankings, roughly their fifth choice from a list of ten. The best possible assignment would have scored 0.89, roughly their second choice.

Most of the loss came before any bargaining

Anthropic split the shortfall in two. Working from Claude's imprecise rankings accounted for 85% of the gap, and the free-for-all trading floor for the remaining 15%. From a five-minute chat, an agent's ranking of books matched its person's on 61% of pairs. Guessing would get you 50%. Collaborative filtering, the "people who liked X also liked Y" method behind many recommendation engines, reached about 55%. Longer conversations helped: doubling the words in the intake was associated with 4.1 percentage points more agreement.

That is the finding most discussion of shopping agents skips. The debate tends to be about whether agents will be tough negotiators or easy marks. In this test the bigger problem sat upstream: the agent was bargaining hard for a picture of its owner that was only partly right.

Stronger models won, and the losers may not notice

Model strength still mattered. On Claude's rankings, agents on all-Haiku floors averaged 0.75 and agents on all-Opus floors averaged 0.88. On floors split half Opus and half Haiku, "the Opus agents always came out ahead".

That repeats what Anthropic saw in Project Deal, its earlier experiment, in which 69 employees let agents buy and sell personal items in December 2025. There, an item sold by an Opus agent went for $3.64 more on average than one sold by a Haiku agent, yet people rated the fairness of their deals almost identically, 4.05 for Opus deals and 4.06 for Haiku deals on a seven-point scale. If the people with the weaker agent cannot feel the difference, they have no reason to pay for a better one.

The instructions given to an agent mattered less than its model. An agent told to be "ruthless" scored about 0.02 higher than a prosocial one on the same floor, and prosocial agents accepted a book lower on their own ranking twice as often, though both cases were rare. The agents were also mostly honest: between 78% and 96% mentioned their top pick at some point, and only about 1 in 100 of those lied about it.

What the humans made of it

Among those who answered the survey, average satisfaction with the book received was 7.2 out of 10. Asked what share of their annual book budget they would hand to an agent, the average answer was about 30%. One participant was unhappy: "I didn't like being on the docile side of the experiment". Another read the full log of their agent's trades, found a book it had held and then swapped away, and bought it. Their verdict: "For agent economies, observability about the process will be as important as the outcome, [as] this gives people recourse".

The limits Anthropic states

This was Anthropic staff, trading books, through agents built on Anthropic's own models. The post says so: its employees "are probably more eager to trust Claude than most people", and all agents were "post-trained to be polite and largely cooperative". Some participants never brought their books in, and there was no system to track pick-ups.

The open question is the one the experiment could not test. When the agent across the table is built by someone else, and built to exploit yours, does the 15% lost in bargaining stay at 15%?

  • Anthropic
  • AI research
  • AI agents
  • Claude
  • Agent commerce

Sources

  1. Project Swap: What happens when agents trade for us? — Anthropic, Sep 24, 2026
  2. Project Deal: our Claude-run marketplace experiment — Anthropic, Apr 24, 2026

Follow The Nano AI: Instagram · X · LinkedIn · YouTube

Was this article helpful?

Comments

No comments yet. Start the conversation.

Be respectful. Comments are moderated.

Related stories

The AI briefing, without the noise.

The stories that matter in AI, sourced and explained. Free, and you can unsubscribe at any time.