
#
Anthropic describes Project Swap as a controlled barter experiment. Participants brought one book they were willing to trade. Claude-powered agents then helped them look for another book.
The study involved 201 employees. They were spread across six pools linked to Anthropic offices in San Francisco, New York City, London, Seattle, Washington, D.C., and Dublin. Anthropic says Dublin was excluded from most analyses.
Participants first discussed their tastes with Claude. The model then estimated a ranking of books in each local pool. Agents negotiated on a digital trading floor. Participants separately ranked ten books so researchers could compare those rankings with the model's estimates. The agents did not see those ground-truth rankings.
Anthropic reports that Claude's pairwise ordering of book preferences agreed with participant rankings 61 percent of the time. That was better than random guessing at 50 percent. It also beat popularity-based guessing at 53 percent and a collaborative-filtering baseline at about 55 percent.
The average outcome in the decentralized market scored 0.55 on a normalized preference scale. The best possible assignment using participants' own rankings scored 0.89. Anthropic says most of that gap came from the model's imperfect estimates of what people wanted, not from the market design itself.
In reruns, stronger models performed better according to Claude's rankings. Anthropic says that did not remove the preference-estimation problem. The study therefore points to a limit in how well the agents could infer human preferences from the available setup.
Anthropic also compared agents instructed to act ruthlessly with agents given an additional prosocial goal. It reports small differences in scores. It also says prosocial agents sometimes made sacrifices.
The source does not provide enough detail to say which instruction was better in every case. It only shows that changing the agent goal altered behavior in limited ways. That makes the comparison useful as an early signal, not a final answer.
Anthropic frames Project Swap as an early investigation into agent behavior, market rules, and participant oversight. It is a small workplace study using books. It is not a deployed marketplace.
The source also says it is not proof that agents can safely negotiate medical, employment, or financial transactions. That distinction matters. A book trade is a narrow test of preference matching and negotiation. It does not establish safety in higher-stakes settings.
The biggest issue in the report is the gap between model estimates and participant rankings. Anthropic attributes most of the performance loss to that estimation problem. In other words, the agents could negotiate, but they still had trouble understanding what people actually wanted.
That finding suggests a practical limit for agent systems that depend on preference inference. If the model misreads the user, the market can still function and yet produce weaker outcomes. The study shows that better negotiation alone may not solve that problem.
The source reports no Morocco-specific facts. For readers, the global lesson is that agent systems can improve coordination only when they estimate human preferences well.
Project Swap is a controlled test of AI agents in a simple exchange setting. It shows measurable progress over random and baseline methods. It also shows a clear gap between agent performance and the best possible human-informed outcome.
Anthropic presents the work as an early step, not a finished system. The study is useful because it separates negotiation from preference understanding. That separation may matter as much as the trading mechanism itself.
Add Intelligence Artificielle Maroc as a preferred source to see more of our relevant stories in Google Search.
We build custom AI platforms, SaaS products, intelligent business applications, and automation systems.
This form is for project inquiries, not general questions about artificial intelligence.