Gradient Flow (Newsletter)
Newsletter (Digital)
Welcome to the nexus of innovation and insights in AI, Machine Learning, and Data! Ben Lorica edits the Gradient Flow newsletter. He helps organize the AI Conference, the AI Agent Conference, the Applied AI Summit, while also serving as the Strategic Content Chair for AI at the Linux Foundation. He is the host of the Data Exchange podcast. You can follow him on Linkedin, Mastodon, Reddit, Bluesky, YouTube, or TikTok. This newsletter is produced by Gradient Flow. Source
Actions
Media Outlet details
| Scope | International |
|---|---|
| Language | English |
| Country | United States of America |
|
Similarweb UVM |
Request pricing |
|
Comscore UVM |
Request pricing |
Recent Articles
Search ArticlesI think we are looking for AI risk in the wrong place
The biggest AI risks sit outside the model The most revealing AI failures right now are not stories about models becoming too capable. They are stories about everything around the model. One system received more access than its test environment could contain. Another was trained on material whose acquisition created $1.5 billion in exposure. In a third, people walked away more certain without being any more correct. I recently argued that passing your evals does not mean you are safe.
Your AI should be better on day 500
A while back I wrote about how startups are using reinforcement learning to make agents more reliable. A deeper problem behind that whole trend keeps resurfacing: a model can improve during training, but the moment it’s deployed, learning largely stops. A policy changes, a new edge case shows up, a user corrects the system, and the lesson rarely travels past that one incident. The prompt gets patched, the ticket gets closed, and the same class of mistake eventually comes back.
Your AI passed its evals. That's the problem.
Evals are part of every serious conversation about putting AI into production. Teams define benchmarks, set thresholds, and increasingly run red teams to see how the system holds up against someone actively trying to break it. That combination is reasonably good at telling you whether a model is accurate, reliable, fast enough for production, and resistant to an adversarial attack. It says almost nothing about what actually gets a company into trouble once the system is live.
The Big AI Labs Are Suddenly Competing with Your Own Data
Specialized AI Is Getting Easier to Build Last week I argued that open models will absorb most of the money and compute the world spends on AI. A week later, open weights are even more central to the conversation. Recent releases have made the gap between capable and affordable harder to ignore, and a broad coalition of technology companies is now publicly arguing that open models matter for competition, security, and national sovereignty. But the models themselves are only the starting point.
Here's my uncomfortable bet on OpenAI and Anthropic
Open Models Will Absorb Most of the AI Spend Here is my bet: open models (open weights and open source alike) will end up absorbing most of the money and compute the world spends on AI. The proprietary frontier models get the headlines and the IPO valuations, but developers and AI teams see something different up close. Open models are improving fast, the gap to the proprietary leaders keeps narrowing, and the economics point one way.
they're literally building tools to catch agents cheating
What Startups Taught Me About the Next Layer of AI Infrastructure A little while back I wrote about how teams use reinforcement learning (RL) to make agents reliable. Since then I keep bumping into startups where RL is not a research footnote or a feature buried in the stack. It is central to what they are building. I know of more than 25 at last count and still climbing, and in some cases RL is basically the product.
I changed my mind about how agents use tools
Your CLI Was Built for Humans, Not Agents There’s a friendly debate among developers about how to give AI agents reliable ways to use external tools, data, and services so they can do useful work beyond generating text. One side favors CLIs, or Command-Line Interfaces, where agents run text commands against tools like and internal scripts.
I talked to Google's former AI head about messy data
Like everyone else, I’ve been enjoying the steady improvement in coding agents and the tooling around them, from frameworks and harnesses to evaluation suites. But the more I talk with teams actually deploying agents in enterprises, the more I circle back to plumbing. Agents need data integrations and infrastructure built for them, a theme I explored in a piece on rethinking databases for agents.
What your base model doesn't protect you from
I’ve avoided writing about copyright and AI. Not because it isn’t important, but because it felt like a legal sideshow compared to the engineering and business questions I find more interesting. That’s gotten harder to justify. There is also something mildly funny about naming my podcast The Data Exchange years ago, because I believed data markets would become central to technology, and then mostly ducking this corner of the data conversation.
The single-vendor AI stack is a transitional phase
The Hybrid AI Stack Is Coming for the Pricing Power of OpenAI and Anthropic OpenAI and Anthropic are going public while still capturing much of the money spent on foundation-model usage. But deployment patterns are starting to tell a more complicated story. Companies are building hybrid model portfolios, using proprietary models where convenience, support, and frontier capability matter, while turning to open-weights models where cost, privacy, customization, and deployment control matter more.