Barak Lenz
Is this you? As a journalist, you can create a free Muck Rack account to customize your profile, list your contact preferences, and upload a portfolio of your best work.
Claim your profile
Get in touch with Barak
Contact Barak, search articles and posts on X, monitor coverage, and track replies from one place.
Learn more about Muck RackActions
Is this you?
As a journalist, you can create a free Muck Rack account to customize your profile, list your contact preferences, and upload a portfolio of your best work.Articles
Jamba: A Hybrid Transformer-Mamba Language Model
Papers Published on Mar 28 · Featured in Daily Papers on Apr 1 Abstract We present Jamba, a new base large language model based on a novel hybrid Transformer-Mamba mixture-of-experts (MoE) architecture. Specifically, Jamba interleaves blocks of Transformer and Mamba layers, enjoying the benefits of both model families. MoE is added in some of these layers to increase model capacity while keeping active parameter usage manageable.
PMI-Masking: Principled masking of correlated spans
PMI-Masking: Principled masking of correlated spans Sep 28, 2020 (edited Mar 18, 2021)ICLR 2021 SpotlightReaders: Everyone Keywords: Language modeling, BERT, pointwise mutual information Abstract: Masking tokens uniformly at random constitutes a common flaw in the pretraining of Masked Language Models (MLMs) such as BERT.
Actions
Is this you?
As a journalist, you can create a free Muck Rack account to customize your profile, list your contact preferences, and upload a portfolio of your best work.Get in touch with Barak
Contact Barak, search articles and posts on X, monitor coverage, and track replies from one place.
Learn more about Muck Rack