What changed
Everyone in AI is on cloud nine. We fly by instruments.
A research studio that builds its own AI companies, then brings what survives to yours.
Research, studio, ventures, and client work, running as one loop.
Since 2007 · New York
Four companies. Three studios. Thirty-seven published experiments. Three shutdowns, written up in full. Nothing on the menu that we have not run ourselves.
Cloud9Lab is a research studio that owns the businesses it learns from.
We have been independent since 2007, which means we have watched a few technologies arrive with this much noise around them. Ten times we have taken a question we could not settle from the outside and built something to find out. Seven of those are still running: four companies and three studios. Three are closed. The write-ups for all three are still on this site, and one of them is the most-read thing we have ever published.
Who we are
We stay small on purpose. The people who run the research are the same people running the ventures and sitting in client rooms. Split that into three teams and the feedback dies, and we are back to selling maps of territory we have never crossed.
So the standard is short. Publish the conditions, because a number without its constraints is advertising. Name the owner, because every shipped decision has a person attached to it. Kill it early, because interest is not traction and a pilot is not a customer. And eat first: if it is not good enough for our own accounts, it is not ready for yours.
The Loop
Consultancies sell maps of territory they have never crossed. We built Cloud9Lab so that could never be us.
Research asks what just became possible. The studio builds it small. The ventures run it against real customers and a real payroll. AX carries what survived into your organization, scars documented. What one part learns, every part inherits.
And when a problem does not fit our current answer, it does not get forced. It gets scheduled: next quarter’s research question.
Research
What just became possible, and under what conditions?
Every Tuesday we publish one finding: what we asked, what we measured, when it holds, and when it breaks. Thirty-seven so far. We do not publish to describe what already happened. We publish to decide what to build next.
Studio
Build it small. Ship it fast. Kill it early.
Weeks one and two are the question. If we can answer it without building, we do not build. Weeks three to six put a prototype in front of ten real users, with no launch and no landing page. By week twelve there is one paying customer, or we stop. Interest is not traction.
Ventures
Real customers. Real payroll.
Seven businesses are running now, four companies and three studios. Each one began as a research question we could not answer from the outside. Three others did not make it, and those write-ups are public in the same series as the wins. A test that cannot lose is not a test.
AX
What survived, brought to your organization.
We do not install software. We draw the line where the machine stops and your name begins, then build the gates that hold it. AI projects rarely die at the model. They die at the moment someone has to sign the output and no one has decided who. Every method we bring you has a scar to prove it.
Models will improve. Interfaces will multiply. Access will become nearly universal. The advantage will not come from having intelligence. It will come from knowing where it belongs, what it is allowed to decide, and who signs when it is wrong.
The new scarcity is judgment
When everyone can rent intelligence, advantage moves to whoever decides better. A model can rank options and spot quality. What it cannot do is carry a consequence. That part still needs a name on it.
Intelligence without authority is a demo
An agent becomes useful only when its permissions, its limits, and its owner are written into the organization. Most org charts do not say whose name goes on an automated decision. That gap is where AI projects die, not at the model.
A roadmap is a document
Transformation is a changed Monday morning. If the queue, the handoff, the exception, and the meeting all look the same as last quarter, nothing has happened yet.
These are counts, not projections. We update them quarterly, and we do not quietly remove the ones that closed.
730+
Across nineteen years, for global brands and for teams of just ten people.
70%
Seven of ten still running. The three that closed are published on this site.
37
Each published with the conditions it holds under, and the ones where it breaks.
Research notes
What we learned, tuition included
Little
“What do you still own when a provider changes terms? We tested five APIs against five rights. Two passed three. None passed five. Holds for rented APIs. Breaks on weights you own.”
Data rights
Five APIs, five rights
Partly
“Does a faster model make the work faster? Not proportionally: five times the inference speed bought twice the turn speed. Holds when tool calls sit in the loop. Breaks on pure generation.”
Latency
Twelve providers, one model
Yes
“Can a model you own beat one you rent? On catalog review, yes: 87% accuracy against 77%. Holds in one vertical with stable rules. Breaks when policy edges turn ambiguous.”
Open models
4,000 items, six weeks
Case studies
Every case ends the same way: the numbers, then the part we got wrong.
D2C retail
Three full-time roles spent the week checking product listings, and errors still reached the storefront. Nothing was allowed to leave their own infrastructure.
6h
/ per week
Down from 120 hours, live in eleven weeks.
- Storefront errors down 31%
- Now runs in-house for about $180 a month
- Reviewers moved to the calls that need a person
- Where we were wrong:
the first version silently rejected 30% of valid listings
Freight forwarding
Every shipment needed a tariff code, and the two people who knew the edge cases were both near retirement. A wrong code meant a fine, not just rework.
9min
/ per shipment
Down from fifty minutes, live in nine weeks.
What changed
- Disputed classifications down 22%
- The two-person dependency is gone
- Every code now carries its citation
- Where we were wrong:
we optimized for speed and buried the reasoning
Consumer brand
Regional teams needed far more content, and conventional automation produced work that was fast, consistent, and forgettable. The voice was the asset at risk.
4.6x
/ regional variants
Review hours flat, live in six months.
What changed
- Cost per approved asset down 38%
- Review time per asset down 78%
- Local teams kept the final call
- Where we were wrong:
we automated approval first and had to put a person back
