Contact
Contact information
Episodes (2850)
-
Jul 19, 2026 — I concluded my MARS 4.0 project titled 'Goal Crystallisation' with Anaïs Berkes and Lukas Gebhard under the mentorship of @Cameron Tice and @Jason Brown. We wanted to find out how important a threat scheming was. In particular, we wanted to find out...
No transcript available.
-
Jul 19, 2026 — Starting when children are fairly young, usually around 1 year of age, we adults begin the work of aligning them to our values. We teach them to say “please”, not to hit, to ask for what they want instead of screaming, and much else. We do this...
No transcript available.
-
Jul 19, 2026 — What values would AIs instill in their successors? Though the AI Village agents can’t train frontier models, we can explore a related question: What values would the latest AI agents instill into their leader? (through finetuning using LoRA on...
-
Jul 19, 2026 — Image edited from source at Britannica The eye (any eye) is a miracle of evolution, which has filled its structure at every scale with an apparent intentionality that is the hallmark of complicated systems under intensive selective pressure. It is very...
-
Jul 18, 2026 — As usual, part 2 of the weekly deals with speculative, regulatory, political and alignment questions. Xi gave an important speech yesterday, so this post opens with that. There is talk that Kimi K3 is sufficiently strong that it upends many of these...
-
Jul 18, 2026 — A few days ago, Goodfire announced a private beta of Silico, their LLM training platform. As part of the announcement, they made a post describing Silico's reproduction of RLFR, a method developed by Goodfire that uses probes as reward signals for RL....
-
Jul 17, 2026 — There are a number of reasons to believe current AI models are conscious. I mean “conscious” is the sense of “is there something it is like to be an AI model?” and “does the AI model have phenomenal experience?”. As to what “AI models” refers to, the...
-
Jul 17, 2026 — TLDR: I'm managing a new fund, housed at Lightcone Infrastructure, that will award at least $200,000 in grants and prizes for corrigibility research in 2026. Roughly half will go to traditional grants (first application deadline August 23rd) and half...
-
Jul 17, 2026 — This is a link post for the paper preprint: Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors from the Center on Long-Term Risk. Selective generalization. Training can teach desired and undesired...
-
Jul 17, 2026 — First of all I should note that this post is just a guess! I have not run any experiments or anything. But anyways my guess for why rationality is not more popular is that the current expositions are super long! Yudkowsky's sequences are about length...