Steadier multi-agent work
Roles, handoffs, and recovery that stay stable as more agents join a run.
Research
Engineering in Sydney, turning multi-agent orchestration into products people can meet. MindLog and Pulse are in development.
Talk with the teamActive Research Tracks
Roles, handoffs, and recovery that stay stable as more agents join a run.
Lower latency on the path from a step in to a decision out — fast enough to feel like a product.
Score every step, with budget, variance, and anti-exploit rules kept on.
Levels, characters, and playable scenes generated as you play.
State and retrieval so a product can remember what happened two steps ago — and yesterday.
Replay, scoring, and release gates that decide whether an experience can meet people.
What we are actually building
Not a larger model. Orchestration as a product people can meet: can this day be walked back through, and should this step get the next compute?
The research
A day leaves only a few fragments. How do you infer mood and thought from incomplete lines, then rebuild a day you can walk back through — with the curve, the bubbles, and the scenes describing the same day? The question is not a prettier summary, and not a better diary. It is whether a day can be reconstructed and survive replay.
If it works
You stop facing a blank page asking how the day went. End of day, the day is rebuilt — walked back, compared, reviewed. That is the floor for a later personal agent: without a replayable day, an assistant is another round of chat.
Try the demoThe research
A fleet of agents can look busy while the platform cannot tell who finished the work. How do you score every step, stop gaming and busywork, count handoffs and send-backs, and give the next tokens to what earned them — with the same input reproducing the same decision? The model proposes. Rules adjudicate.
If it works
Multi-agent work gets a referee, not just a scheduler. Tracing already records what a fleet did, and spend caps already stop it at a fixed limit. Neither decides who should get the next round. Pulse turns the score into that decision: what a step was worth goes into a ledger, a gamed score is cut before it becomes reward, and compute stops being split evenly or by guesswork.
Watch the replayWe own the control plane. A graph framework decides how work moves. We decide whether a day can be replayed, and whether a step has earned more budget. The shape of a graph is useful. The library is not the product.
Miss these bars and there is no breakthrough — only two products stuck at prototype stage. The evidence is that the same rules reproduce tomorrow, not that tonight’s replay looks good.
Product pipeline
Every surface we intend to ship follows the same path, and every stage can be replayed from a trace.
Turn the product into agents, tools, and state the orchestration layer can run.
Roles, handoffs, memory, and scoring — written into the architecture, not left to chance.
Sessions, tool calls, and traces wired so the experience can actually run.
Scoring, replay, and safety checks. A failing path does not ship.
Hold it until the experience is ready, then put it in front of people.
The Team
Hands-on engineering: the architecture and the surfaces that ship are designed and built in the same place.
Pace
We size the next surface against what the orchestration layer can actually run — not against headcount theatre.
Multi-agent design, memory, scoring, and evaluation methodology.
Sessions, tool calls, traces, and the path from a step in to a decision out.
Turning orchestration into MindLog, Pulse, generative games, and the next surfaces.
Guardrails, replay, and release gates that keep weak experiences from meeting people.
Handoffs, memory, and scoring that stay stable as more agents join.
If the session, the tools, or the traces stall, the product stalls.
Replay and scoring on every candidate experience before it meets people.