← All articles

Deterministic Multi-Agent Orchestration on MCP

A team of agents needs someone to decide what happens next. The obvious candidate is another model : An orchestrator agent that reads each result and picks the following step. The CompanyManager plugin makes the opposite choice. The sequencing of a run lives in engine code. A model does the work of one step and ends its answer with a verdict ; The engine reads that verdict and decides where the run goes. This article describes how that engine is built.

Sequencing is code, work is the model

The engine is a class named PipelineEngine. It takes the workflow that a team declares, runs its steps one mission at a time, hands the accumulated context to each step, and routes the run on the structured verdict that any step may return. The contract is the same for every step. A review is a step whose instruction asks for a judgement ; The engine has no special case for it.

The graph : Steps and labelled edges

A workflow is a graph with an entry stage, a list of nodes, and a list of edges. A node wraps one step : A stage name, the capability that must carry it, an optional extra instruction, and a flag that places a human gate after it. The node adds a visit budget and, optionally, a project event to emit when the node passes. Editor coordinates are stored with the node and ignored by the engine.

An edge has three fields : The source stage, a verdict label, and the target stage. Two values are reserved. The target __end__ terminates the run. The label * is a catch-all, followed when no other edge of the node matches the verdict.

Nodes hold no routing logic. Every decision about the next step is an edge, so the whole control flow of a team is readable as data and editable without touching the engine. The agent teams page shows the graphs that run in production, drawn from that same data.

Static validation at save time

A graph is validated when it is stored, before any run uses it. The validator returns a list of errors, and a graph with one error is refused. It checks that :

  • The entry is present and names an existing stage
  • Every edge starts from an existing stage and points to an existing stage or to __end__
  • Every edge carries a verdict label
  • No node has two outgoing edges for the same label, compared without regard to case
  • A termination is reachable from the entry : A node without outgoing edges, or an edge to __end__

The fourth rule is what makes routing a function. Given a node and a verdict, at most one edge applies, so the engine never has to choose between two candidates. The fifth rule rejects a graph that could only absorb a run without ever finishing it.

One check produces a warning instead of an error : A step whose capability no agent of the team carries. Such a graph is accepted, and the step is assigned at run time to the team’s full-authority agent until a specialised role exists.

Cycles are allowed, and bounded

The validator does not reject cycles. A cycle is how iterative correction is expressed : A review node with a FAIL edge back to the implementation node is a loop, and it is the intended behaviour. What the engine controls is how many times a loop may turn.

Each node carries a maxVisits budget, three by default. The engine keeps a counter per stage for the run and increments it every time the run enters the node. When the counter exceeds the budget, the run stops and goes to a person, with a reason that names the node and its budget. Oscillation between two steps is therefore a property of the engine, not a behaviour that a prompt has to prevent.

Two kinds of re-run do not consume a visit. A human gate that returns amendments re-runs its node, and the person at the gate is the guard of that loop. A re-ask for a malformed verdict, described below, is free as well.

The verdict contract is derived from the edges

An agent does not need to know the graph. Before each mission, the engine reads the outgoing edges of the current node, builds the list of labels they carry, and appends a short contract to the instruction of the step : End the answer with a line of the form VERDICT: label, where the label is one of that list.

The contract has a single source, a function on the graph, and it adapts to the node. A node whose only exit is PASS receives no contract at all : A missing verdict line is read as PASS. A node with several exits including PASS receives the list and the same default. A node with no PASS edge, a triage step for instance, receives a stricter text : Neither PASS nor FAIL exists here, and an answer without a verdict line cannot be used.

The rule behind this is that a worker is told only the verdicts its node can route. Labels are free text chosen by the team, so a graph can route on business outcomes as easily as on pass and fail. Routing itself stays in code : The engine looks for the edge whose label equals the verdict, then falls back to the catch-all edge if the node has one.

Context is passed in cascade, then compacted

A run owns one context string. It starts with the client’s request. After each step, the engine appends the full result of that step under a header that names the stage, and the next mission receives the whole context. A step therefore sees what every previous step produced, without any agent calling another.

Appending without limit would let old results dominate the prompt of every later mission. Right after the append, the engine compacts the context in tiers, from the most recent result to the oldest :

  • The preamble stays intact : The client’s request and the answers a person gave at a gate
  • The most recent results stay intact
  • The results before them are cut down to a fixed byte budget
  • The oldest results keep their header only

Routing lines are always kept : Verdicts, amendments, and gate validations are short and carry the thread of the decisions. The compaction is idempotent, which matters because a resumed run goes through it again at every step. Nothing is lost by it either ; The full result of each step remains in the record of its mission.

Escalation : When the engine stops and asks

The engine never improvises a route. Three situations stop a run and hand it to a person.

  • A node has exhausted its visit budget. The correction loop does not converge, and another turn would not change that.
  • A verdict matches no outgoing edge. The engine asks the same node once more, with a note that lists the labels the node routes and asks for the verdict line only. If the second answer is still not routable, the run escalates.
  • A mission ends in a status other than done. The run is recorded as failed at that stage, with the cause taken from the output of the worker.

In each case the engine stores the reason with the run and posts a message that names the stage and reports what the worker said. Two reserved verdicts exist outside any graph. One states that inputs are missing : The run moves to awaiting_input and the questions of the worker go to the person. The other states that an upstream project lacks a capability : The run moves to awaiting_upstream with the identifier of the feature request that the worker opened, and it resumes when that request is delivered.

A human gate is the planned form of the same idea. When a node carries the gate flag, the engine records the run as gate_wait and blocks before following any edge. An approval lets the run continue, and the annotations given with it enter the context. A refusal adds the requested amendments to the context and re-runs the node.

A run is a journal, so it can resume

The engine writes the state of a run at every edge transition, at the opening of a gate, and at every end. The record holds the graph name, the status, the current stage, the visit counters, the cascaded context, the recap, the original request, the gate in progress, and the git branch dedicated to the run. The record is not the engine ; It is its trace.

That trace is what makes a stop cheap. A run in status failed, escalated, interrupted, gate_wait, awaiting_input, or awaiting_upstream is resumable, including after a restart of the server. Resuming re-enters the graph loop at the current node with the context restored and the visit budgets re-armed. Exhausted budgets are a common reason for a stop, and a resume is an explicit decision by a person, so the person is the one who re-arms them.

A stop request follows the same path. The engine checks for it at the top of each iteration and again when a mission returns. It records the run as interrupted at the current node, which is resumable as is. A deliberate stop is not reported as a failure.

One engine thread per run, in the server process

Starting a run returns to the client at once. The engine launches a detached thread inside the MCP server process, and that thread executes the graph loop from the entry node to the end. Each mission is awaited in that thread : The loop starts a mission, blocks until it returns its status and result, then routes. A human gate blocks the same thread on a condition variable until the person answers.

The sequencing therefore needs nothing outside the process that already hosts the plugins. State that must survive the process is the run journal described above. The thread catches every exception, reports it to the conversation, releases the lock that marks the team as busy, and replays the next queued trigger for that team.

Where this fits

The engine sits above two mechanisms described elsewhere. Each mission it starts is an instance of a typed template, covered in Mission-as-Template : Declarative AI Agents in Production. Each worker runs the reason-act loop dissected in Building a ReAct Agent on Top of MCP. The engine adds the layer between missions : Which step comes next, how many times, and who is asked when the answer is not in the graph. For the roles, the capabilities, and the supervisor and worker split that the graph relies on, see Company Manager ; For the graphs deployed today, see agent teams.

Talk to us about MCP infrastructure for your organisation. Get in touch.