Key takeaways from {unscripted} San Francisco's panels on AI agent trust, security, and adoption. Plus, where {unscripted} is headed next.

TL;DR
I've sat through a lot of conference afternoons in my time, and last Thursday at Oracle Park was one where I actually kept my phone in my pocket. {unscripted} 2026 kicked off its roadshow in San Francisco, and the afternoon sessions turned into one of the most candid conversations I've been part of on AI in the software delivery lifecycle. Three back-to-back panels on trust, security, and adoption surfaced themes I can't stop thinking about. Here's my recap, plus where to catch the show this week.
Writing code that runs inside your heart
The CIO of a global medical device company opened the afternoon with a reminder of what's actually at stake when "move fast" meets literal life-or-death software. Her company treats roughly 75 million patients a year (two every second) across cardiac, spine, diabetes, and surgical robotics. Some of that code runs inside a patient's body. I don't think I'll forget that framing anytime soon.
Her point, and it's one I've made in different words for years: regulatory and compliance work isn't a brake on velocity, it's the discipline that lets you safely go faster. Because product cycles in medical devices run in years, not days, the goal of AI tooling here isn't shipping faster for its own sake, it's giving engineers better visibility into risk, security, and regulatory implications earlier in the process, across the 170 countries her company operates in.
Speed vs. trust
In a panel of experts, moderated by one of our own from the Harness team alongside leaders from a frontier AI lab, a data platform company, and an AI customer-experience company, tackled the false dichotomy of the day head-on: is speed or trust more important when deploying AI agents?
The consensus that emerged was more nuanced than either/or, and honestly, it's the same conclusion I keep landing on:
- Trust is the floor, not the ceiling. Nobody actually knows where AI capability tops out right now, so chasing the ceiling is a waste of energy. Trust and quality bars should stay high regardless of whether work is pre- or post-AI. I don't care what era of tooling you're in, the bar doesn't move.
- The real trust question is about guardrails, not the agent itself. One panelist put it well: think about onboarding a new engineer. You don't watch every keystroke, but catastrophic mistakes are structurally out of reach for them. Agents need the same structural constraints, "but enhanced, because they like to go fast, and they don't sleep."
- Reversibility is the deciding factor. If an action is easily reversible, let it move fast. If it's hard to reverse, add friction. As one panelist memorably put it, agents can behave like "a caffeinated 12-year-old who knows how to code." You want the guardrails of a responsible adult in the room. I've said for years that governance isn't the thing slowing you down; the lack of it is what eventually does.
- "Context native" means more than RAG. Several panelists pushed back on the idea that bolting a retrieval system onto an agent counts as context. The real gap is in what one called "the graveyard" of undocumented Slack threads, tribal knowledge, decisions that live only in one person's memory. This is exactly the "library without a card catalog" problem I talk about all the time. The information exists, but finding the right piece at the right moment depends on who happens to remember it.
- The panel closed on hot takes, and one landed harder than the rest: "My personal hot take is if you have a roadmap, you're not moving fast enough." The room reacted like it was a hot take, but nobody from product got up to disagree with it either. I'm still chewing on that one.
Security teams are losing the head start
The security panel with engineering leaders from an automation software vendor and a global medtech company opened with a stark stat: the average time for an attacker to weaponize a disclosed vulnerability has dropped from 125 days to about a day. Defenders, meanwhile, are still averaging roughly seven days to patch. That gap is the whole ballgame.
Here are some key takeaways:
- Frontier models are currently better at precision than recall on vulnerability detection. They're good at confirming what they find, less good at finding everything that's there.
- Models marketed as "cyber-specific" aren't meaningfully outperforming general-purpose frontier models on these benchmarks yet.
- Scan results are still probabilistic: the same scan on the same repo can return different findings, take different time, and cost differently run to run.
- The structural fix isn't a separate security review gate. It's collapsing security into the same pull-request-level workflow as everything else, so vulnerability detection happens at the speed code actually ships, not on a separate cadence. This is the same argument I make about governance generally: it only works when it's embedded in the pipeline, not bolted on beside it.
What actually breaks when a whole org adopts AI
The final panel, featuring engineering leaders from a major workforce platform, a consumer review and local-business platform, and others, got into the operational reality of scaling AI coding org-wide. It's where the Golden Canary story came full circle for me.
One leader shared what happened after a routine 10% increase to their AI coding tool licenses: fifteen days into the new licensing year, they'd blown through 700 new license requests in a month. The cause wasn't just more engineers using the tool, it was marketing, support, and customer success teams all picking it up too, alongside agents generating code around the clock. Code volume roughly quadrupled. I've seen this movie before, just with a different tool cast in the lead role.
That exposed the real bottleneck: code review capacity hadn't scaled with code generation, and some teams generating code didn't have engineers available to review it at all. The description of unmanaged AI-generated code at scale that's going to stick with me: "a room full of kindergartners on their first day with one teacher trying to manage it." Standing up code-review agents helped, but it also generated its own new review load, meaning the whole system, not just the generation step, needs to scale together. Fast deployment doesn't mean much if recovery, or in this case, review, can't keep pace.
The closing thread tied back to the earlier panels: as AI shifts developer time away from typing code and toward writing and refining specs, the bottleneck moves with it. One session noted that roughly 80% of developer time is now going into spec-writing and review, not coding. That tracks with everything else I heard on that stage.
What's coming this week
San Francisco was just the first stop. {unscripted} 2026 continues as a roadshow, and three more cities are up this week:
- Chicago — September 15
- Boston — September 16
- New York City — September 17
Columbus and Paris follow on September 22, London and Dallas on September 24, Atlanta on September 29, and a virtual {unscripted} wraps the series on September 30. If the SF sessions are any indication, expect more of the same hard questions: how much autonomy to give agents, how to keep security in the loop instead of behind it, and how to scale review capacity as fast as generation capacity grows. I'll be at a few of these myself, so come find me if you want to debate about trust versus speed in person. I've got opinions.



%20(1).jpg)