Is this you? As a journalist, you can create a free Muck Rack account to customize your profile, list your contact preferences, and upload a portfolio of your best work.
Claim your profile
Get in touch with Chitta
Contact Chitta, search articles and posts on X, monitor coverage, and track replies from one place.
Learn more about Muck RackActions
Is this you?
As a journalist, you can create a free Muck Rack account to customize your profile, list your contact preferences, and upload a portfolio of your best work.Articles
Structured reasoning failures compromise LLM interpretation of clinical oncology notes
Abstract Large language models (LLMs) show strong performance on clinical benchmarks, yet their reasoning reliability in real-world oncology care remains unclear. We evaluated LLM reasoning on authentic oncology notes using a novel hierarchical error taxonomy across two retrospective cohorts spanning breast, pancreatic, and prostate cancer. GPT-4 produced reasoning errors in 23.1% of note interpretations, the majority reflecting cognitive bias patterns.
[2512.15949] The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
Actions
Is this you?
As a journalist, you can create a free Muck Rack account to customize your profile, list your contact preferences, and upload a portfolio of your best work.Get in touch with Chitta
Contact Chitta, search articles and posts on X, monitor coverage, and track replies from one place.
Learn more about Muck Rack