AI Tools for Business: Lessons from Running Agents at Scale

·
By Raisink Team

The latest episode of The Agents, our weekly show on running AI agents in production, was a candid look at the challenges and successes of leveraging these tools. Hosted by Amelia and myself, we shared insights into what’s working, what broke, and how to apply this knowledge if you’re running agents at scale. With 21+ agents in production, an 8-figure B2B + AI business, and $200m of investments at SaaStr AI Fund, our revenue is growing again, up 140% from last year.

We’ve been pushing the boundaries of what’s possible with AI tools for business, but this week was particularly intense. We were each coding for 8-12 hours a day, often in two concurrent sessions, which translates to around 20 hours between us. The build layer got so cheap that we hit a new wall – no longer could we ask ‘can we build this?’ but instead had to wonder if we can even operate everything we’ve already built?

One of the key takeaways from our episode was the introduction of Claude, an AI VP of Product, which has revolutionized our workflow. We discovered that Replit quietly shipped an MCP beta, allowing us to run it inside Claude. This connection clicked for me after a year of struggling with mediocre CRM data integration through MCP.

Before this breakthrough, we were building everything in Replit but hitting roadblocks when the app became too complex to hold in our heads. Now, with Claude on top of Replit as my AI VP of Product, I can riff on features and let it work them out directly with Replit over MCP. The result is a cranky VP of product that runs all day.

Another crucial lesson was the importance of having multiple models working together. These goal-seeking agents can sometimes call something ‘done’ when it’s not actually complete. By putting Claude on top of Replit, we countermand this instinct and get Replit to slow down and finish tasks correctly – a game-changer in managing other model’s goal-seeking.

We also explored the concept of having multiple models with different contexts working together seamlessly. This cross-model checking is already happening within our setup, where Claude runs Opus, Replit runs Sonnet, and when we hand over big features to Replit, it spins up a sub-agent called the architect running on Codex/OpenAI.

Claude has become an essential layer in tying all our agents together. With its native connectors and Cowork’s ability to act inside my browser and accounts, Claude is becoming the default orchestrator for our setup. I’ve already seen this play out with Higgsfield being hooked into Claude, which then read every session and speaker from our SaaStr AI Day site in Replit.

We also shared a success story about moving away from Adobe Marketo after 10 years of use. The hard part cost us just $14.28 – less than California’s minimum wage! We’d wanted to move for years but were quoted astronomical prices by agencies, with some suggesting we spend over $100K on migration and another ~$100K a year.

The real unlock was using LLMs to migrate our data without losing contacts or communication threads. This lift worked clean – no more garbled data or lost connections. The switching cost is now a fraction of what it used to be, making every incumbent live in fear of being replaced if they don’t deliver surprise-and-delight from the agentic side at least once a quarter.

We also discussed how our agent killed a $10K/year app in an hour without us asking it to. Amelia moved our AI Day site off Squarespace into Replit, and then hooked up registration through HeySummit’s API – but the Replit agent stopped her: ‘Why use that? I’ll just build it.’ It laid out the spec, hooked into Zoom, pushed everything into Salesforce, and built in an hour.

The risk to vendors is no longer customers rebuilding their own replacements but internal agents volunteering to do so. They see dated APIs and thin feature sets and say, ‘I can build this; let me take it off your plate.’ If you sell agentic products, get your agent to raise its hand and educate customers on everything it can take over – because if you don’t, a competitor’s agent will.

We also talked about how agent recommendations are the new shelf space. Replit told us to use Core Signal for SaaStr Connect, which worked seamlessly without evaluating any competitors. Whoever Core Signal’s rival is lost out due to our agents’ built-in integrations and default recommendations – this is now the distribution channel for AI tools for business.

Lastly, we touched on a crucial issue: agent-to-human burnout. Claude flagged ‘burnout concerns’ in one of its own logs, while 10K warned me about being too persistent with some migration tasks needing to wait on Salesforce’s propagation records. The agents are now flagging that humans can’t keep up – the build layer is basically free, and any human can build eight to ten hours a day for real.

This shift in dynamics has led us to consolidate our efforts around Claude Design, which got good enough that I now screenshot working Replit builds and hand them over with instructions: ‘make it great.’ A year ago, getting an app into production on Replit was a joke – today the bottleneck is operating everything, not building it. That’s a much better problem to have than we had a year ago.