Building AgentTalkie
One conversation across a growing fleet of agents
Agents are multiplying faster than anyone can supervise them. AgentTalkie is one conversation across all of them.
I built it at the AI Tinkerers Global Hackathon: a voice workspace for your agents. Talk through the work, approve an action, and see the tool call, result, and receipt. In the recorded run, AgentTalkie created a real Ambiguous document and read it back to prove the write completed.
I run Grokbot, OpenClaw, Hermes, Claude Code, and Codex, and the list keeps growing. Cloud and local, each with its own session and memory. There is no central place to coordinate the fleet or unblock the work that needs my decision.
Agents work well by themselves. The fleet is not coordinated.
What I built
The prototype directs four kinds of work from one conversation:
- Documents: Ambiguous AI drafts a document, pauses for approval, validates against the live schema, writes it, and reads it back.
- Research: Exa returns source-backed results with the query, endpoint, latency, and cost visible.
- Coding: Ori launches Codex on the local machine and streams its native output into the workspace.
- Handoff: Claude Code receives the exact artifact Codex produced, edits it, and returns a new playable version without replacing the last good result.
Voice is the interface. Durable jobs, approvals, event streams, and artifact identity are the system underneath it.
GPT-Live is not Realtime
GPT-Live uses a separate surface: POST /v1/live/sessions, gpt-live-1, client delegation, and WebRTC. Existing Realtime code does not simply port over.
Client delegation was the unlock. The voice model does not need to answer from its own context. It can hand work to the right agent, keep the conversation moving, and surface the result when that work finishes.
Long-running work still has to leave the audio turn. If a job takes 40 seconds, the interface needs an explicit running state and progress. Silence in a live conversation feels like failure.
Receipts make the system believable
CopilotKit and AG-UI provide a typed event stream, so the action trail is made of real events rather than parsed logs. A proposal, approval, tool call, result, and receipt are distinct states.
That mattered most with Ambiguous AI. Its 17 applications sit behind one MCP surface and are designed around agent actions. The best implementation decision of the day was simple: validate against the live schema before a write, then re-read the result afterward.
The readback is not decoration. It proves the system changed the intended object instead of merely saying it did.
Latency is a product decision
Exa profiles behave like a latency dial. Instant is roughly 250 milliseconds, fast roughly 450 milliseconds, and deep reasoning can take around 40 seconds.
One search in the recorded workflow returned 10 results in 1,006 milliseconds for $0.007. The query, endpoint, sources, timing, and cost were visible. That is fast enough to stay inside a voice interaction. A 40-second search belongs in a background job with progress.
The right model is not always the most capable one. It is the one whose latency and cost fit the moment in the workflow.
Local agents change the trust boundary
OpenRouter’s Ori was the surprise. It does not run a hosted worker for this flow. It spawns codex exec as a child process on the operator’s machine and uses the installed Codex authentication.
Our run inherited six local MCP servers and the local skill catalog. That makes Ori powerful because the coding agent can work with the operator’s real environment. It also means isolation is opt-in, not the default. Before handing it a job, the operator should be able to see the model, budget, workspace, available tools, and execution boundary.
Routing remained explicit through OpenRouter. Local execution did not mean invisible model selection or unbounded spend.
What I learned
One conversation is not enough by itself. Coordination requires durable state and visible evidence.
- Voice should delegate work, not pretend to complete every task inside the audio turn.
- Writes need approval, schema validation, and readback.
- Research needs sources, latency, and cost attached to the result.
- Agent handoffs need exact artifact identity, not a summary of what the previous agent made.
- Local execution needs an explicit trust boundary.
- Every action should leave a receipt.
One very fun day hacking away at a prototype: voice, tasks, research, and a coding worker. Each with a receipt.
Try AgentTalkie or read the project case study.