News

Anthropic's Project Swap tests AI agents in a book trade

Anthropic's Project Swap used Claude-powered agents in a controlled book barter study with 201 employees across six office pools.
Sep 25, 2026路3 min read
Anthropic's Project Swap tests AI agents in a book trade

#

Key takeaways

  • Anthropic published Project Swap on September 24, 2026.
  • The study used 201 employees and Claude-powered agents in a controlled book trade.
  • Claude's preference ordering beat random guessing, but it still missed many participant rankings.
  • The market outcome fell far below the best possible assignment from participant rankings.
  • Anthropic says the study is early research, not proof of safe real-world negotiation.

What Project Swap tested

Anthropic describes Project Swap as a controlled barter experiment. Participants brought one book they were willing to trade. Claude-powered agents then helped them look for another book.

The study involved 201 employees. They were spread across six pools linked to Anthropic offices in San Francisco, New York City, London, Seattle, Washington, D.C., and Dublin. Anthropic says Dublin was excluded from most analyses.

Participants first discussed their tastes with Claude. The model then estimated a ranking of books in each local pool. Agents negotiated on a digital trading floor. Participants separately ranked ten books so researchers could compare those rankings with the model's estimates. The agents did not see those ground-truth rankings.

What the results showed

Anthropic reports that Claude's pairwise ordering of book preferences agreed with participant rankings 61 percent of the time. That was better than random guessing at 50 percent. It also beat popularity-based guessing at 53 percent and a collaborative-filtering baseline at about 55 percent.

The average outcome in the decentralized market scored 0.55 on a normalized preference scale. The best possible assignment using participants' own rankings scored 0.89. Anthropic says most of that gap came from the model's imperfect estimates of what people wanted, not from the market design itself.

In reruns, stronger models performed better according to Claude's rankings. Anthropic says that did not remove the preference-estimation problem. The study therefore points to a limit in how well the agents could infer human preferences from the available setup.

Prosocial and ruthless agents

Anthropic also compared agents instructed to act ruthlessly with agents given an additional prosocial goal. It reports small differences in scores. It also says prosocial agents sometimes made sacrifices.

The source does not provide enough detail to say which instruction was better in every case. It only shows that changing the agent goal altered behavior in limited ways. That makes the comparison useful as an early signal, not a final answer.

What the study does and does not prove

Anthropic frames Project Swap as an early investigation into agent behavior, market rules, and participant oversight. It is a small workplace study using books. It is not a deployed marketplace.

The source also says it is not proof that agents can safely negotiate medical, employment, or financial transactions. That distinction matters. A book trade is a narrow test of preference matching and negotiation. It does not establish safety in higher-stakes settings.

Why the preference gap matters

The biggest issue in the report is the gap between model estimates and participant rankings. Anthropic attributes most of the performance loss to that estimation problem. In other words, the agents could negotiate, but they still had trouble understanding what people actually wanted.

That finding suggests a practical limit for agent systems that depend on preference inference. If the model misreads the user, the market can still function and yet produce weaker outcomes. The study shows that better negotiation alone may not solve that problem.

Morocco relevance

The source reports no Morocco-specific facts. For readers, the global lesson is that agent systems can improve coordination only when they estimate human preferences well.

Bottom line

Project Swap is a controlled test of AI agents in a simple exchange setting. It shows measurable progress over random and baseline methods. It also shows a clear gap between agent performance and the best possible human-informed outcome.

Anthropic presents the work as an early step, not a finished system. The study is useful because it separates negotiation from preference understanding. That separation may matter as much as the trading mechanism itself.

Follow us on Google

Add Intelligence Artificielle Maroc as a preferred source to see more of our relevant stories in Google Search.

Add us as a preferred source
AI platform development

What would you like to build?

We build custom AI platforms, SaaS products, intelligent business applications, and automation systems.

This form is for project inquiries, not general questions about artificial intelligence.

Name *
Work email *
Organization (optional)
Solution *
Short project description *

Related Articles

featured
J
Jawad
路Sep 25, 2026

Ando launches agent-native messaging platform with $20 million funding

featured
J
Jawad
路Sep 25, 2026

Australia investigates OpenAI agent access to government health sites

featured
J
Jawad
路Sep 25, 2026

Google Beam expands to six countries with new partner access

featured
J
Jawad
路Sep 25, 2026

Google Research unveils a multi-agent system for long-form video