Gradient Flow (Newsletter)
Newsletter (Digital)
Welcome to the nexus of innovation and insights in AI, Machine Learning, and Data! Ben Lorica edits the Gradient Flow newsletter. He helps organize the AI Conference, the AI Agent Conference, the Applied AI Summit, while also serving as the Strategic Content Chair for AI at the Linux Foundation. He is the host of the Data Exchange podcast. You can follow him on Linkedin, Mastodon, Reddit, Bluesky, YouTube, or TikTok. This newsletter is produced by Gradient Flow. Source
Actions
Media Outlet details
| Scope | International |
|---|---|
| Language | English |
| Country | United States of America |
|
Similarweb UVM |
Request pricing |
|
Comscore UVM |
Request pricing |
Recent Articles
Search ArticlesThe part of the AI security story I keep thinking about
If you check Ethics.Dev regularly, you may have noticed how often AI agents and cybersecurity have been colliding lately. Some of the stories sound straight out of movies: agents escaping test environments, communicating through channels nobody intended, finding vulnerabilities, stealing credentials, and reaching production systems. The part I think deserves more attention is less dramatic. AI is changing the economics of an attack. A human attacker has limited time and attention.
Your model is easy to buy. Your data pipeline is not.
Last year I wrote that robotics had a data problem language models were lucky enough to avoid. Text, images, and code had accumulated for decades before anyone decided to train models on them. Robots had no comparable internet of physical experience. That distinction is starting to look less clean. Recent systems have trained on enormous collections of human video. One company reported using more than a million hours.
Your AI might be fine but your data is lying to it
Stale Data Used to Be Annoying Imagine a procurement agent checking inventory and deciding that stock has fallen below the reorder threshold. It places another order. The problem is that a large delivery was recorded a few minutes earlier, and the copy of the data the agent queried has not caught up. A dashboard running on stale data might show you the wrong inventory number. An agent running on stale data can place the wrong order.
What counts as valuable company data is changing
When Work History Becomes an AI Asset Google recently agreed to pay $10 million for the internal business data of Spirit Airlines. The airline is bankrupt, but its emails, Teams messages, software, spreadsheets, and operating records apparently still have value. Google plans to use the material for product development and AI.
This is how self improving AI actually starts
A few weeks ago I wrote about AI systems that keep learning after deployment instead of treating every interaction as a fresh start. That prompted a few readers to ask about recursive self-improvement (RSI). The connection is real, but I think it helps to separate three ideas. Continual learning asks whether experience makes a system better next time. Bounded self-improvement goes further by letting the system help turn that experience into a tested, persistent change.
I think AI teams are defending the wrong thing
Your Model Is Not Your Moat In a recent article, I wrote about companies turning their own data, workflows, and production feedback into specialized intelligence they increasingly control. The deeper idea was compounding. What matters is not simply whether a model performs well today, but whether using it creates assets that make the system better tomorrow. More AI teams are beginning to build around that idea.
I keep hearing the same advice about agents
Nine Practical Rules for Agents Doing Real Work In recent conversations with crews building agents, I keep hearing the same lessons. Teams with very different products are arriving independently at almost the same architectural choices. That convergence feels important. In recent posts, I argued that passing your evals does not mean an AI system is safe, and that many of the most consequential risks sit outside the model itself.
I think we are looking for AI risk in the wrong place
The biggest AI risks sit outside the model The most revealing AI failures right now are not stories about models becoming too capable. They are stories about everything around the model. One system received more access than its test environment could contain. Another was trained on material whose acquisition created $1.5 billion in exposure. In a third, people walked away more certain without being any more correct. I recently argued that passing your evals does not mean you are safe.
Your AI should be better on day 500
A while back I wrote about how startups are using reinforcement learning to make agents more reliable. A deeper problem behind that whole trend keeps resurfacing: a model can improve during training, but the moment it’s deployed, learning largely stops. A policy changes, a new edge case shows up, a user corrects the system, and the lesson rarely travels past that one incident. The prompt gets patched, the ticket gets closed, and the same class of mistake eventually comes back.
Your AI passed its evals. That's the problem.
Evals are part of every serious conversation about putting AI into production. Teams define benchmarks, set thresholds, and increasingly run red teams to see how the system holds up against someone actively trying to break it. That combination is reasonably good at telling you whether a model is accurate, reliable, fast enough for production, and resistant to an adversarial attack. It says almost nothing about what actually gets a company into trouble once the system is live.