TheSequence
Newsletter (Digital)
The best source to stay up-to-date with the developments in the machine learning, artificial intelligence, and data science world. Trusted by 165,000 professionals from the main AI labs, universities, and enterprises Source
Actions
Media Outlet details
| Scope | National |
|---|---|
| Language | English |
| Country | United States of America |
|
Similarweb UVM |
Request pricing |
|
Comscore UVM |
Request pricing |
| Frequency | Other |
Recent Articles
Search ArticlesThe Sequence AI of the Week #899: Inside Inkling: A Trillion-Parameter Model That Only Wakes Up 41 Billion at a Time
Inkling is best understood not as a single 975-billion-parameter brain that fires all at once, but as a giant warehouse of specialist capacity. A router chooses a small working set for each token. That makes the arithmetic sparse, while storage, networking, and deployment remain very large. The headline number for Inkling is 975 billion parameters. That is close enough to a trillion that the distinction is mostly useful to accountants.
The Sequence Knowledge #898: The Trace Is the Teacher: Distilling Reasoning Into Small Models
In January 2025, DeepSeek took its big reasoning model, R1, and used it to generate around 800,000 worked solutions — long, rambling chains of thought, complete with false starts, self-corrections, and the occasional “wait, let me reconsider.” They filtered these traces for correctness and readability, and then did the most boring thing imaginable with them: plain supervised fine-tuning on a handful of off-the-shelf open models — Qwen at 1.5B, 7B, 14B, 32B; Llama at 8B and 70B.
The Sequence Radar #897: Last Week in AI: China, Compression and the Open-Model Race
We continue our series about model distillation techniques. In the AI of the Week , we discuss Thinking Machine first open weights model. The opinion section reviews the idea that Google’s is by far the biggest threat to NVIDIA’s dominance. TheSequence is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. For years, AI progress has been narrated as a horse race: larger models, higher benchmark scores, more expensive clusters.
The Sequence Opinion #896: Spark, Compute, and the Two Metas
Last Thursday, Mark Zuckerberg posted on X for the first time in three years. That alone should tell you something. The occasion was the launch of Muse Spark 1.1, the second model out of Meta Superintelligence Labs and the first Meta model ever to ship with a price tag. It arrived with a public API, aggressive pricing at $1.25 per million input tokens and $4.25 per million output tokens, an OpenAI-compatible endpoint, and closed weights. Read that last part again.
The Sequence AI of the Week #895: OpenAI's Show Us Where Coding Evals Break
OpenAI’s audit of SWE-Bench Pro shows why a precise score can still be a poor measure - and why coding agents may become essential tools for auditing the benchmarks that grade them. A frontier coding score can look wonderfully precise: 80.3 percent, one decimal place, clean enough to rank models and anchor product claims. But precision is not validity.
The Sequence Knowledge #894: When the Student Started Talking Back: Distillation in the LLM Era
Looking back at the 2015 distillation paper, what’s striking isn’t the temperature trick or the dark-knowledge framing --- it’s the world the paper quietly assumed. There was a fixed input distribution. There was a teacher that produced a probability vector over a closed set of classes. There was a student trained to match that vector. Run the dataset through both, compute the loss, backprop. The pipeline had a kind of mechanical innocence to it. Everything stayed in its lane.
The Sequence Radar #893: Last Week in AI: GPT-5.6, Grok 4.5, Muse Spark 1.1 and the Post-Chatbot Stack
We continue our series about model distillation. In the AI of the Week, we are going to discuss OpenAI’s recent analysis of coding benchamarks. In the opinion section, we are going to debate Meta’s opportunities and tremendous challenges to catch up with the AI frontier labs. TheSequence is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
The Sequence Opinion #892: The Anatomy of a Good Environment: When Verifiability is Not Enough
Quick note: For over two years now, we’ve been running The Sequence without sponsors. Literally every week we receive tens of inquires to sponsor The Sequence. As a passion project that I have been running for over five years now, we are trying to keep The Sequence true to its ethos of technical depth, objectivity and unique content to help you navigate the world of AI. The best way to help us on that journey is to subscribe below. We understand is not for everyone but appreciate the support.
The Sequence AI of the Week #891: Prompting a Spreadsheet : Inside Google’s TabFM for Tabular AI
Quick note: For over two years now, we’ve been running The Sequence without sponsors. Literally every week we receive tens of inquires to sponsor The Sequence. As a passion project that I have been running for over five years now, we are trying to keep The Sequence true to its ethos of technical depth, objectivity and unique content to help you navigate the world of AI. The best way to help us on that journey is to subscribe below. We understand is not for everyone but appreciate the support.
The Sequence Knowledge #890: A Brief History of Model Distillation
The story most people tell about knowledge distillation starts in 2015, with Geoffrey Hinton, Oriol Vinyals, and Jeff Dean introducing a clever softmax temperature trick and a phrase — “dark knowledge” — that immediately lodged itself in the field’s vocabulary. It is a good story. It is also incomplete by almost a decade.