Articles/Field notes

Field notes

Multi-Agent has a Complexity Problem


Multi-Agent AI Has a Complexity Problem

At AI on the Amstel’s HumanX kickoff in Amsterdam, practitioners pushed back against the race to build increasingly elaborate agent systems. Their message: start with one agent, prove where it fails, and only then add complexity.

Multi-agent AI sounds compelling.

A network of specialized agents coordinating with each other, handing off work and eventually operating like an autonomous organization.

But at the AI on the Amstel kickoff meetup for HumanX Amsterdam, the discussion was considerably less futuristic — and arguably more useful.

Moderated by Grant Easterbrook, the panel brought together Vladislav Tankov, Anthony Diaz and Sean Kenny to discuss what multi-agent AI looks like once you move beyond demos and start deploying it inside real organizations.

The recurring message was simple:

Most companies should probably use fewer agents than they think.

Start with one agent

The panel repeatedly challenged the assumption that a sophisticated AI system needs to begin with multiple agents.

Instead, the suggested progression was much closer to conventional software engineering:

Start with one agent.

Understand where it succeeds and where it consistently fails.

Then isolate the part of the workflow causing the problem and, if necessary, give that responsibility to another agent.

Sean Kenny described an account-management use case where a single agent initially handled the whole workflow: understanding an account, analyzing data and preparing a report.

Only once the team saw that data analysis was becoming a bottleneck did they separate that task into a dedicated sub-agent.

The distinction matters.

The second agent was not added because a multi-agent architecture looked more advanced. It was added because there was a measurable limitation in the simpler system.

One line from the discussion captured the broader philosophy:

“Complexity looks cool and sophisticated on management slides. Avoid it at all costs.”

The burden of proof should therefore be reversed.

Do not ask why you should keep the system simple.

Ask why the additional agent needs to exist.

Multi-agent AI is not a token-saving strategy

Another misconception addressed during the session was cost.

Running several agents does not inherently make a workflow more efficient.

In many cases, it does the opposite.

Each agent has to understand its environment, load context, reason about its task, communicate and potentially validate the work of other agents.

That can create enormous duplication.

Kenny described inspecting agent traces and discovering that a main agent would first understand the environment and formulate a plan, only for the sub-agent receiving the task to repeat much of the exact same process.

The optimization did not come from adding another agent.

It came from reading what the agents were actually doing.

That led to one of the most practical recommendations of the panel:

Read the transcripts.

Agent-to-agent communication can look elegant from the outside while underneath, agents may be repeating context, interrupting each other or performing unnecessary reasoning.

The panel's argument was not that parallel agents are inherently bad.

They can be extremely valuable when speed, specialization or parallel execution justifies the additional compute.

But the metric should be the cost of achieving a useful outcome — including retries, duplicated work and human review — rather than simply counting tokens.

Observability becomes non-negotiable

As agent systems expand, another problem appears:

Understanding what is actually happening.

Once five, ten or potentially hundreds of agents are executing tasks, organizations need to know:

  • What each agent is responsible for
  • Which tools it can access
  • What it costs
  • Whether its outputs meet expectations
  • What happens when something goes wrong

That makes observability less of an infrastructure nice-to-have and more of a prerequisite for deploying agents at scale.

The panel repeatedly returned to the need for evaluation, logging, access controls and clear responsibilities.

This becomes especially important because adding agents can amplify problems rather than solve them.

If the original agent is failing because it lacks context, receives poor instructions or uses an unreliable tool, creating five additional agents can simply reproduce the same failure five times.

The advice was closer to debugging a distributed system than managing an imaginary workforce:

Understand the failure first.

The real AI rollout problem is organizational

Perhaps the most important part of the conversation had very little to do with models.

Building or buying the technology, the panel argued, is often the easy part.

Getting an organization to actually use it is harder.

Engineering teams can provide infrastructure, but they do not necessarily understand the detailed workflows of legal, logistics, sales, finance or marketing teams.

Domain experts do.

The result is that successful enterprise AI adoption cannot simply be delegated to a central AI team.

Companies need people inside individual departments who understand both the work and enough of the technology to redesign how that work gets done.

Kenny described these people almost as AI architects embedded inside teams.

And the panel was skeptical of traditional corporate training as the solution.

A mandatory course explaining how to use AI is very different from giving employees access to the technology and letting them discover what it can change in their own workflow.

Managers using the tools themselves, hands-on experimentation and internal power users appeared repeatedly as more credible mechanisms for adoption.

The underlying message was significant:

AI transformation is increasingly an organizational design problem, not simply a software deployment problem.

Budget: control it — but don't kill experimentation

Here the panel showed more disagreement.

One view emphasized establishing budgets and monitoring from the beginning.

Without visibility into AI expenditure, organizations can suddenly discover millions being spent and respond by shutting experimentation down.

Another view argued almost the opposite:

Companies sometimes need to accept a temporary “stupid tax” while employees discover valuable use cases.

Premature optimization can prevent those use cases from ever being found.

Both arguments point toward the same tension enterprises now face.

AI experimentation has to be constrained enough that it does not become financially uncontrolled, but open enough that employees are able to discover workflows management could never have designed centrally.

There is unlikely to be a universal answer.

What matters is knowing whether increased spending is producing increased capability.

Don't marry your stack too early

The same caution applied to technology providers.

Models, agent frameworks and orchestration systems are changing too quickly for today's best-performing stack to be assumed to remain the best choice several years from now.

That creates a particular problem for enterprises accustomed to multi-year software agreements.

The panel discussed two possible strategies.

One is to remain deliberately flexible:

Preserve the ability to switch models, mix providers and use specialized agents where they perform best.

The other is to commit deeply to one ecosystem and use the commercial benefits, integrations and discounts that come with that relationship.

Where the speakers differed was on which strategy companies should choose.

Where they largely agreed was that accidentally becoming locked in is the worst of both worlds.

Security is still mostly a software problem

When the conversation moved toward AI swarms and security, the panel pushed back against some of the more dramatic narratives around thousands of autonomous agents suddenly operating beyond control.

For most companies, the immediate problems are more familiar:

  • Access control
  • Permissions
  • Data exposure
  • Monitoring
  • Tool usage
  • Budget limits
  • Human approval for consequential actions

In other words, much of AI security still resembles traditional software and infrastructure security.

Agents should only receive the permissions they actually need.

Sensitive actions should require additional controls.

External content should be treated carefully.

And organizations need the ability to inspect what happened and stop a run when necessary.

The panel also argued that the bigger near-term cybersecurity concern may not be multi-agent coordination itself, but increasingly capable models becoming better at cybersecurity tasks in general.

Small models still matter

During the Q&A, the discussion also turned to smaller and open models.

The argument was straightforward:

Not every task requires the most expensive frontier model.

For narrowly defined work, smaller models can often be faster, cheaper and easier to evaluate.

This becomes particularly interesting in multi-agent systems, where a powerful central agent might delegate clearly defined subtasks to much smaller models.

The principle again comes back to specialization.

If a task is predictable and well-scoped, the most powerful model available may simply be unnecessary.

AI organizations may eventually look like human organizations — or not

Despite spending much of the session criticizing unnecessary multi-agent complexity, the panel was not dismissive of multi-agent AI itself.

Quite the opposite.

Several speakers expect systems of cooperating agents to become considerably more important over the next few years.

Companies already divide work between departments, specialists and managers.

It is intuitive that agent systems may eventually replicate parts of that structure:

Specialized agents executing tasks, other agents coordinating them and humans interacting with progressively higher layers of abstraction.

But there was also a counterpoint.

AI systems do not necessarily have to reproduce organizations designed around human limitations.

The eventual architecture may look less like a company made of artificial employees and more like something humans would never have designed in the first place.

That may be the more interesting prediction.

We may end up using increasingly powerful systems while understanding less about why their internal organization works.

The takeaway

The conversation around AI agents has moved very quickly from chatbots to assistants, agents, teams and now swarms.

But terminology can move considerably faster than production systems.

The strongest lesson from AI on the Amstel was therefore not anti-agent or anti-multi-agent.

It was more disciplined:

Start simple. Define the outcome. Evaluate it. Inspect what the system actually does. Add specialization only when you can explain why it is necessary.

A thousand agents make a compelling demo.

One agent reliably solving a real business problem may still be more valuable.


Panel

  • Grant Easterbrook — Moderator
  • Vladislav Tankov
  • Anthony Diaz
  • Sean Kenny

Event

AI on the Amstel — HumanX Amsterdam kickoff

Amsterdam, September 2026.

People involved

Grant Easterbook

Moderator · AI On the Amstel

LinkedIn

Anthony Diaz

Panelist · Orq.ai

LinkedIn

Sean Kenny

Panelist · Firmament

LinkedIn

Vladislav Tankow

Panelist · JetBrains

LinkedIn