Yingzhe Peng
Is this you? As a journalist, you can create a free Muck Rack account to customize your profile, list your contact preferences, and upload a portfolio of your best work.
Claim your profile
Get in touch with Yingzhe
Contact Yingzhe, search articles and posts on X, monitor coverage, and track replies from one place.
Learn more about Muck RackActions
Is this you?
As a journalist, you can create a free Muck Rack account to customize your profile, list your contact preferences, and upload a portfolio of your best work.Articles
Paper page - LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL
TL;DR: We present a two-stage rule-based reinforcement learning approach that significantly improves multimodal reasoning capabilities in LMMs, enabling our 3B model to outperform much larger models like Claude 3.5 and GPT-4o on agent tasks.
Actions
Is this you?
As a journalist, you can create a free Muck Rack account to customize your profile, list your contact preferences, and upload a portfolio of your best work.Get in touch with Yingzhe
Contact Yingzhe, search articles and posts on X, monitor coverage, and track replies from one place.
Learn more about Muck Rack