Heading to Las Vegas this week for Ai4 at The Venetian 🎲
A conference about non-deterministic systems, in a city built entirely on the fact that you can't predict the outcome!
I can't wait to connect, learn, and share with AI leaders shaping what comes next. If you're there, come find me. I'd love to talk about how we're thinking about Context & Agent Trust at Monte Carlo.
Here's the idea I keep coming back to: we spent decades perfecting the Software Development Life Cycle (SDLC), how about Agent Development Life Cycle (ADLC)?
The two rhyme, but the differences are exactly where the hard problems (and the 2am pages) live.
Learnings thus far:
1️⃣ Evals are the new QA
Evals in agent development are the equivalent of QA testing in software. A great foundation. But like QA, you can't predict every way real users (and agents) will interact with your system, or the beautifully cursed questions they'll actually ask. Green evals and a live incident can absolutely coexist!
2️⃣ The Golden Dataset is controlled. Context is a group chat!
Golden Dataset that sat still and behaved but Agents run on context: retrieved, dynamic, and changing while you're looking at it. You don't own that input surface anymore
3️⃣ Humans could observe and fix. Agents laugh at that plan.
In traditional software, a human could observe, detect, and resolve issues. With agents, the complexity, dependencies, and sheer volume make manual oversight roughly as scalable as counting cards with your eyes closed. Trust has to be engineered in: observable, measurable, and continuous (Reinforcement Loops?)
Software taught us that shipping is the beginning, not the end. Same for agents, except now the ground moves too.
If this resonates, or you're quietly living it in your own stack, let's connect. I'll be at Ai4 all week, statistically somewhere near the coffee or at Monte Carlo Booth # 1260 👋
And if you like racing, join us Wednesday from 7 to 10 PM.
#Ai4 #AIAgents #AIObservability #MonteCarlo #AgentDevelopment #DataTrust
👏