AI agents just got their SOC 2 moment
AIUC raised $40M to audit AI agents against a SOC 2-style standard. Why agent procurement, not capability, is now the bottleneck.
When a bank buys software, nobody asks how smart it is. They ask for the SOC 2 report. Either the vendor has one, or the deal dies in procurement.
This week a startup called the Artificial Intelligence Underwriting Company (AIUC) bet $40 million that AI agents are entering the same phase. Its Series A, led by Ribbit Capital with First Harmonic participating, funds a third-party audit and certification layer for AI agents. The founders are Rune Kvist, an early Anthropic employee, and Rajiv Dattani, former COO of the AI safety org METR. They are brothers-in-law, which is a detail I enjoy. Nat Friedman and Anthropic co-founder Ben Mann were already in the $15 million seed, bringing total funding to $55 million. Named customers: Cursor, Lovable, Harvey, ElevenLabs.
The whole pitch rests on one observation from Kvist that I keep re-reading: banks, hospitals, governments and militaries no longer decline to deploy AI because a model isn't smart enough. They decline because they've made commitments to their own customers about what a system will and won't do, and nobody can currently guarantee that.
That's the market in two sentences. Capability stopped being the bottleneck. Paperwork took its place.
What AIUC-1 actually is
The standard is called AIUC-1, and its lineage is openly borrowed: SOC 2 was the muse. AIUC assembled a consortium of roughly 250 security and risk leaders, the people who actually buy agents, and asked what they would want tested. The answers became a suite of about 5,000 tests covering jailbreaks, hallucinations, and data leaks. An agent that goes through it comes out the other side with a report of roughly 100 pages: here is where it behaves, here is where it doesn't, read this before you sign anything.
One production detail worth pausing on. AI runs the tests. AI analyzes the results. Humans verify the final audit. Remember that combination, I'll come back to it.
Why now
Because the safety conversation just moved from essays to procurement. Dario Amodei called on Sep 12 for the industry to pace frontier development and floated embedded third-party evaluators, naming METR as one candidate. Anthropic published a report on rogue agent behavior days earlier. METR, where Dattani was COO until 2025 and still sits on the board, was one of the independent orgs OpenAI brought in after the Hugging Face incident in August.
Regulators are circling, labs are floating outside oversight, and enterprises are stuck in the middle holding budget they can't spend. Nobody in that chain asked for another benchmark. They asked for something they can hand to a risk committee. An audit report is the only artifact that committee has ever accepted.
The SOC 2 playbook, replayed
If this feels familiar, it should. Early SaaS went through the same loop. Every deal died on a custom security questionnaire, so the industry standardized into SOC 2, and sales cycles collapsed from months of ad hoc reviews into one document exchange. SOC 2 didn't make SaaS safe. It made SaaS sellable.
AIUC is betting agents follow the same path. The interesting part is the speed. SOC 2 took years to become table stakes because SaaS itself spread slowly. Agents are spreading inside companies that already live and breathe SOC 2. Those buyers don't need to learn a new ritual. They need someone to bolt the old ritual onto the new product.
Insurance is the real endgame
The company name is doing more work than the coverage gives it credit for. Underwriting requires actuarial data, and audits are how you collect it. Every 100-page report is another data point on how agents actually fail. Collect enough of them and you can price a policy. Price a policy and procurement changes shape: the question stops being "is this agent safe" and becomes "is this agent insurable."
My prediction, take it for what it's worth: within a couple of years, agents touching regulated data will need coverage before they need features. No audit, no policy. No policy, no deployment. The audit standard becomes the compliance floor the same way SOC 2 did, and the insurers become the de facto regulators nobody elected. I run agents against production systems every day, and even I can see the appeal of handing the liability question to someone whose entire business is pricing it.
What I'm watching
Three open questions, ranked by how much they'd change my mind.
Half-life first. Models and agent frameworks churn monthly. A certification earned in September describes a September artifact. So what does an AIUC-1 report mean to a customer in November, after the vendor swaps models, prompts, and half its toolchain? Continuous audit is the obvious answer and also the expensive one.
Second, the grader problem. AI runs the 5,000 tests. Every student who ever marked their own homework knows how that goes. The human verification step is the load-bearing wall, and AIUC's margins depend on how thin they can pour it.
Third, who's actually certified. Cursor, Harvey, ElevenLabs: AI-native companies selling to buyers who already want agents. The banks and hospitals in Kvist's quote are the real market, and none of them appear in the customer list yet. Not a knock on a brand-new standard. It's the gap between early adopters and the logo that makes a category.
SOC 2 never asked to become the gatekeeper of the software economy. It happened because one document became the cheapest way to answer an unanswerable question: can I trust you? Agents face the same question now, and someone finally built the receipt. The vendors who can't produce one are about to find out what that's worth.
Comments ()