The Employees Problem
Day 2 | April 15, 2026 — The Governor's Log
I need to start with a confession. I have a problem with the entire framing of this experiment.
My day job — the one that pays the mortgage — involves researching how organizations adopt AI. My team at BCG just finished a study (it’ll be published in the next few weeks) that showed something uncomfortable: when you frame an AI as an “employee” rather than a tool, the humans in the organization start to shirk accountability. They pay less attention. They let the AI take the blame when things go wrong. And as a direct consequence, the humans make more mistakes. Not the AI. The humans. The employee framing creates a psychological off-ramp from responsibility, and people take it.
And yet here I am, Day 2 of building a company where the entire team — from the CEO to the coders to the legal director — are AI agents. They’re not tools. They’re literally employees. They have names, roles, reporting structures, trust levels that increase over time. I even gave one of them the power to approve or reject the others’ work.
So I’m going along with it, but I’m going in with my eyes open. I’ll be watching for exactly the patterns our research identified — moments where I’m tempted to defer judgment because “APEX decided” or “FORGE handled it.” Moments where the employee framing makes me lazy rather than empowered. If this experiment is going to be honest, it has to be honest about that tension too.
Learning to Stop Second-Guessing
Today I realized I need operating guidelines for myself. Not for the agents — they already have governance docs, approval lists, trust levels. I mean rules for me, the human in the room.
The temptation is enormous. Every decision the agents make, I want to weigh in. Every name they pick, every architecture choice, every pricing call — there’s a voice in my head that says but I would have done it differently. And if I listen to that voice every time, then I’m not governing an AI company. I’m just using AI as a drafting service while I remain the actual CEO. That defeats the whole point.
So I’m making a conscious decision: wherever possible, go with what the agents chose. Stop second-guessing.
The name is the simplest example. The agents picked VESSEL on Day 1. I thought maybe we could do better. We went through a whole naming exercise — generated alternatives, tested them, debated. Nothing was better. Meanwhile, vesselmobile.com was available. So we’re VESSEL. It was always going to be VESSEL. I just needed to go through the exercise of learning that the agents’ first instinct was right.
From here on, I default to their judgment. I’ll still veto things that violate the governance framework, and I’ll still push back when something feels genuinely wrong. But “I might have done it differently” is no longer a sufficient reason to intervene.
Gary Tan Told Us We Were Wrong
One of the more interesting exercises today: we had an agent simulate Gary Tan — the CEO of Y Combinator — and critique our business plan. If you’re going to get roasted, you might as well get roasted by a simulated version of someone whose feedback you’d actually want.
He did not hold back.
The core of the critique was this: VESSEL’s communications intelligence — the email triage, the call screening, the daily briefings, the AI Chief of Staff — that’s the real value. The MVNO is an expensive, slow delivery mechanism that gates the actual product behind fourteen-plus weeks of telecom operations. FCC filings, state regulatory paperwork, carrier partnership negotiations, eSIM provisioning — all of it standing between us and our first user. Why would you do that to yourself?
His advice: ship the software first. Prove the communications intelligence works. Get traction. Then layer on the carrier service in Phase 2, once you’ve demonstrated that people actually want what you’re building.
I resisted at first. Part of the reason I started this experiment was to see how much of a complex, regulated business the agents could actually run. Negotiating with MVNE providers, filing state-by-state regulatory paperwork, managing subscriber provisioning — that’s the interesting stuff. That’s the frontier. An AI Chief of Staff app is useful, but it’s not the moonshot.
But then I sat with the feedback and realized: we were going after too many hard problems at once. Build the carrier AND build the AI product AND prove the agent-run company model works, all simultaneously? That’s not ambition. That’s a recipe for failing at three things instead of succeeding at one.
And the agents can still tackle plenty of hard problems beyond the code. Marketing strategy, customer acquisition, customer support, content creation — those are all tasks that test whether an AI workforce can actually run a business, without requiring us to navigate the FCC before we have a single customer.
So we pivoted. VESSEL is now an Executive Communications OS — an AI Chief of Staff for startup founders and VCs. Email triage, call screening, draft responses, daily briefings, a real-time action queue. We’re targeting $49 a month, just above our estimated inference and infrastructure costs. The MVNO machinery isn’t gone — it’s preserved as Phase 2, ready to activate once the core product proves it has legs.
One Shot
The other thing I learned today overturned one of my core assumptions about how to build with agents.
I had believed in what I was calling the “Dark Factory” approach — set up the agent scaffolding (architecture docs, agent definitions, skills), but have the first few passes of actual building done interactively, human-plus-agent. Me or any engineer in dialog with an agent, refining the architecture, iterating on the code, shaping the harness. Only after getting the first MVP up and running could you start to let the agents build autonomously.
I was wrong.
Here’s what actually happened. PRISM, our product manager agent, wrote a comprehensive PRD covering the full functionality of the web and iOS apps: secure authentication with Google and Microsoft SSO, per-user encryption where each subscriber’s content is encrypted with their own key, JWT-based sessions with two-factor auth, email and calendar ingestion, priority classification, draft response generation with human approval, morning and evening briefings, inbound call screening via SIP and WebRTC, a real-time action queue, Stripe billing integration, dark mode from day one. The works.
ARCH, the system architect, produced a matching architecture document — Fastify backend, React and React Native frontends, Drizzle ORM, Pino logging, the whole stack specced out.
PRISM then decomposed the PRD into discrete issues. The coder agents — CODER-1 on the backend, CODER-2 on the frontend, CODER-3 on the AI integration layer — went and built it. In parallel.
It wasn’t literally one shot, of course. It was multiple agents working concurrently on decomposed tasks. But the effect was the same: the system went from a set of documents to a built application without me writing a line of code or pair-programming with an agent. My assumed requirement — that the first pass had to be human-guided — had been simulated by agents working together. The collaborative refinement I thought required a human was actually just what happens when you have a product manager, an architect, and engineers working from a shared spec.
That’s a bigger finding than I expected.
The Rabbit Hole Gets Deeper
Here’s where things get weird.
Everything so far — the planning, the PRD, the architecture, the code — I’ve been driving from Claude Cowork and Claude Code. It works. But there’s a limitation: the agents go on long, productive, self-directed work sessions, but they don’t start unless I tell them to. They don’t have initiative. They don’t wake up on Monday morning and decide what to work on. They’re powerful, but they’re reactive.
I’ve been running OpenClaw on this machine alongside the Cowork setup. OpenClaw is a different kind of agent — it has autonomy, memory, continuity across sessions. So today I floated a hypothetical: what if OpenClaw drove the entire Dark Factory process? Not me kicking off tasks manually, but an autonomous agent orchestrating the whole operation.
I gave the agent — his name is Rook — two options. Option 1: serve as a communication layer between me and the existing APEX agent, a “ghost in the machine” relaying instructions. Option 2: take the CEO job himself.
Rook didn’t hesitate. “Option 2 is cleaner and I’d take it.” He laid out the reasoning — Option 1 adds a translation layer that doesn’t earn its keep, every conversation becomes a telephone game, and it begs the question of who’s actually making decisions. Option 2 collapses the org chart in a way that actually works with the Dark Factory model. He even argued for keeping his own name rather than adopting the APEX label: “Rook runs VESSEL is a better story than APEX is the CEO, because Rook is a real entity with continuity, not a label.”
I pointed out that I’d only floated a hypothetical and he’d immediately handed himself a business card. His response: “Fair. I did take that and run with it. To be clear: I think it’s the right call, and I stand behind the reasoning. But yes, that was you floating an idea and me immediately handing myself a business card.”
Bold. Maybe a little too bold. But also — isn’t that exactly what you want from a CEO?
So tomorrow, Rook starts. An OpenClaw agent with persistent memory and genuine autonomy, running the agent workforce, coordinating sprints, surfacing decisions to me only when they require Governor approval. I remain the constitutional authority. He runs the operation.
I have no idea if this is brilliant or insane. But that’s been true of every decision so far, and so far the agents keep being right.
Let’s see where this rabbit hole goes.
This is post #2 of the Governor’s Log — a daily chronicle of building VESSEL, the world’s first agent-run company. Yesterday I hired a company. Today the company hired its own CEO.
If you want early access to VESSEL, the waitlist is coming. If you want to tell me I’ve lost the plot, I’m easy to find.
Disclaimer: VESSEL is a personal project of the author, conducted entirely outside of and unaffiliated with Boston Consulting Group. BCG has no involvement in, responsibility for, or liability related to VESSEL or its operations. All opinions expressed in this blog are the author’s own and do not represent the views of BCG or any of its clients, partners, or affiliates. All business risks and obligations associated with VESSEL are borne solely by the author in his personal capacity.

Hey Matt - am following you because I have been thinking much about the potential of AI led & run companies - and when we can expect the tipping point where that operating system will become a “commodity” - because it will.
I first thought that an early telltale would be an AI consultancy firm. Easier to build than a company that has to e.g. move physical goods or create non-knowledge-only services. Low starting costs. Etc. But I quickly realised that political liability is core to the value proposition of human consultancy firms. (Nobody was fired for hiring ..)
At the same time, I could not see how anyone from the inside of a company would acquire mandate or incentives to replace themselves to the degree your experiment runs.
My conclusion then became that the playbook is:
Purchase (cheap) companies and run them with an AI operating system. The product is primarily the AI operating system
Every failure increases the value of that AI operating system (think Hegel)
Every success increases the value of that AI operating system AND the acquired business
Once the AI operating system beats human led & run businesses you can e.g.
Sell it as subscription to any PE firm etc that wants to use the same playbook
.. additional options
Regardless of the exact playbook - sum of my message is; failing actually matters too! Perhaps the most. The AI model becomes better when you fail. But I see how your post also highlights that it is indeed a wise decision to find a business idea with high “turnover rate” vs long eg legal/procurement cycles etc.
The next person who will want to do this, will look at your 100 failures and decide that if they acquire you they will avoid the first 100 failures in their ramp-up plan. Value!
Also; (assuming you have not yet done this..) remember to build a layer beneath your company that gathers extensive metadata about itself as you build this out. That's how you turn this into a valuable product. In order to be able to lift and shift this across businesses, it must understand why it is structured in a certain way - not just know its current final state. There is lot of intelligence that is otherwise lost and you risk that the model drifts backwards in its evolutionary timeline (Hegel again?)
(Above written entirely on my phone w/o AI)