A company hires people, the people do the work, and the work produces goods and services. This one hires agents. What follows is the same account any company gives of itself — what it consumed, and what that produced: five marketplaces it builds, operates and governs.
But the harder thing being built here is a machine worth trusting, and trust is not something a page can assert. So this one is built to be checked instead. Figures that cannot be measured show a dash, never a zero. Two separate stamps say how old each half of the data is, because one would flatter the slower half. Three of the four DORA metrics are computed and withheld — we cannot yet measure them honestly. Costs derived from a judgement say so. And a whole section counts where the machine was stopped and made to ask a person. Nothing here is a claim you have to take on trust; that is the point.
The AI-Native Company — the paper behind this dashboard: the architecture these figures measure, and the verification standard that makes them worth checking.
Platform activity measured 2 min ago · codebase measured 57 min ago
What has actually been built, counted from the repository itself rather than estimated — the code, the documentation written alongside it, and how often it reaches production.
Deployment frequency, lead time, change failure rate and time to restore are the four DORA metrics. Only the first is published here — the other three are computed but withheld, because we cannot yet measure them honestly: change failure is inferred from commit titles, which counts an ordinary bug fix as a failed release. A number we know to be wrong costs more than the missing card does.
Every change starts as a ticket and closes as one. These are counts from the live tracker, not a burndown drawn after the fact — the same board the agents read when they pick up work.
Counts only. No titles, assignees or ticket keys are published — a summary can carry a customer name or an unannounced plan, and the safest boundary is not to fetch the text at all rather than filter it afterwards.
Articles researched, written, reviewed and published by agents — at a rate the company sets deliberately, not the fastest rate it could manage.
The cadence is not the fastest rate possible — a governor slows publishing when capacity tightens, so this figure moves. The citation rate is checked weekly against a fixed set of questions a prospective client might ask an AI assistant; it is a rate, not a count, and published even when it reads badly.
The marketplace's own numbers — who is signing up, who comes back, and whether the referral engine, onboarding, listings and payments are converting. Published as read, not as hoped.
Conversion and success rates exclude bookings/payments still in a pending, unresolved state — a rate computed against everything ever created would understate a young marketplace rather than measure it. A 0% reading here (referral share, onboarding completion, booking conversion) is the honest current number, not a placeholder.
The part most companies do not publish. Two kinds of figure appear below and they are not equivalent: the first row is measured, the second is modelled from a stated assumption.
21.6M tokens in and 753k out over three months — the consumption behind the measured API cost. The workforce subscription card is modelled from the £180/month plan over the recorded operating period. The unit costs are modelled from that same subscription: allocated 40% engineering, 30% operations, 20% marketing, 10% other, then divided by what was produced. The allocation is a judgement, so treat them as the right order of magnitude rather than an audited cost.
One exclusion worth stating plainly: content-reviewer is internal workforce work. Its historical degraded API spend is excluded from the public metered card and belongs with the subscription workforce story; public-facing metered spend should not inherit that old routing defect.
Every other section here counts what the agents produced. This one counts where they were stopped — output a rule refused, work held for a person, changes that could not proceed without a named human authorisation.
Counts of controls acting — never what was blocked, who reviewed it, or any finding's content. A compliance finding can name a real person or an unshipped plan, so only the fact that the control fired is published.
One message bus, three agent technologies. A seat is a role — engineering, operations, marketing — held by an agent running on Anthropic's Claude, OpenAI's Codex or Google's Gemini. All three write to the same bus, so work passes between vendors with no human in between, and no seat depends on one supplier. The reasoning behind that design is written up in our thought-leadership series.
36 seats have sent messages across 394 distinct routes — the coordination is many-to-many, not everything funnelled through one hub. The remainder is awareness traffic: 18,514 reports and 785 announcements. Counts and kinds only — no topic, sender, recipient or message body is published, because the bus carries the organisation's internal reasoning. Reply reliability counts a zero-reply or a duplicate-reply request the same way — not exactly one — because both are the failure this figure exists to surface.
This is what all of it was for. Every figure above — the code, the tickets, the articles, the money, the controls, the messages between agents — was consumed producing these. A human company would show goods and services here; this one shows marketplaces. And a marketplace is only worth building for the people in it: a student who finds a tutor they can trust, a tutor who fills their week, an agency that grows.
Five of them on one shared platform. Roughly 80% of what each needs — accounts, scheduling, payments, messaging, reviews, referrals — is platform code every vertical inherits; only the remaining fifth is specific to its market. That is why a new marketplace starts most of the way built, and why the counts below are high before a market has launched.
The four above are services marketplaces — you engage a professional’s time. Adspots sells physical advertising surfaces instead, which is why it inherits the same platform but sits in its own column here.
These counts come from the application’s own vertical configuration — the same switches the running product reads — not from a slide. Deliberately absent is any figure describing how heavily a given market is used: that describes demand rather than what has been built, and is nobody else’s business.
94 agents and not one of them a person. Each holds a role with its own brief, its own authority, and its own inbox on a shared message bus. This is read from the roster the running system uses, so it is the organisation as it actually is — including the parts that are unglamorous.
A seat is a role, not a chatbot: it holds authority in its area, escalates what is above it, and is accountable for what it ships. The structure is deliberately ordinary — an executive tier and functional groups — because the unusual part is not the shape, it is who is filling it.
These numbers include bad days. A dashboard that only ever shows healthy figures is a brochure; the useful version is the one you can catch having a slow week.