Amazon Science
Blog,
Online/Digital
The Amazon Science website gives insight into the company’s approach to customer-obsessed scientific innovation. Amazon fundamentally believes that scientific innovation is essential to being the most customer-centric company in the world. It's the company’s ability to have an impact at scale that allows us to attract some of the brightest minds in artificial intelligence, machine learning, and related fields. Source
Actions
Media Outlet details
| Scope | International |
|---|---|
| Language | English |
| Country | United States of America |
|
Similarweb UVM |
Request pricing |
|
Comscore UVM |
Request pricing |
Recent Articles
Search ArticlesSOP-Bench: A new benchmark for evaluating AI agents on real business procedures
A standard operating procedure, or SOP, is the written set of steps an organization follows to correctly complete an important piece of routine work the same way every time. Almost every industry runs on SOPs. A hospital uses one to register a new patient, a logistics team uses one to decide whether a shipment qualifies as hazardous, a bank uses one to verify a new business customer, and a trust and safety team uses one to decide whether to remove a piece of content.
CaliQEC: In-situ qubit calibration for surface code quantum error correction
Xiang Fang, Keyi Yin, Yuchen Zhu, Jixuan Ruan, Dean Tullsen, Zhiding Liang, Andrew Sornborger, Ang Li, Travis S. Humble, Yufei Ding, Yunong Shi Copy BibTeX Quantum Error Correction (QEC) is essential for fault-tolerant, large-scale quantum computation. However, error drift in qubits undermines QEC performance during long computations, necessitating frequent calibration. Conventional calibration methods disrupt quantum states, requiring system downtime and rendering in situ calibration impractical.
Rohith Nama
LLM-based agents struggle to execute complex, multi-step Standard Operating Procedures (SOPs) that are fundamental to industrial automation. Existing benchmarks fail to capture the procedural complexity and tool orchestration demands of real-world workflows. We introduce SOP-Bench, a benchmark of 2,000+ tasks from human expert-authored SOPs across 12 business domains (healthcare, logistics, finance, content
Nandi Subhrangshu
In conversational AI assistants, SLU models are part of a complex pipeline composed of several modules working in harmony. Hence, an update to the SLU model needs to ensure improvements not only in the model specific metrics but also in the overall conversational assistant. Specifically, the impact on user interaction quality metrics must be factored in, while integrating interactions with distal modules
Pairwise ranking outperforms single-action RL for offline explanation selection: A practical lesson
We report a practical lesson from building a GPU-free explainable-recommendation serving stack: explanations are pre-generated offline into a per-item candidate pool, and a small CPU-resident model selects one at request time.
Improving precision of A/B experiments using trigger intensity
Online randomized controlled experiments (A/B tests) measure causal changes in industry. While these experiments use incremental changes to minimize disruption, they often yield statistically insignificant results due to low signal-to-noise ratios. Precision improvement (or reducing standard error) traditionally focuses on trigger observations - where treatment and control outputs differ. Though effective, detecting all triggers (full knowledge) is prohibitively expensive.
CIGE: An agentic AI test case standard
Test automation is moving from fixed script replay to agent driven execution and judgement. An agent may read a goal, inspect product state, call tools, recover from small changes, and collect evidence before reporting a result. Existing test case formats do not give enough structure for this style of execution. Plain prompts mix setup, purpose, rules, and steps. Traditional scripts are repeatable, but they break when the product flow changes.
Forecasting with factor-augmented time series foundation models
Despite their strong zero-shot forecasting capabilities, Time Series Foundation Models (TSFMs) lack mechanisms for incorporating the structured domain knowledge that practitioners need for interpretability and forecast control. Dynamic Factor Models (DFMs) provide this structure, decomposing series into interpretable, adjustable factors like trend and seasonality while capturing shared dynamics across related series.
Improving the lifetime of aluminum-based superconducting qubits through atomic layer etching and deposition
We present a dry surface treatment combining atomic layer etching and deposition (ALE and ALD) to mitigate dielectric loss in fully fabricated superconducting quantum devices. The treatment starts by conformally removing the native metal oxide and fabrication residues from the exposed surfaces through ALE before in situ encapsulating the metal surfaces with a thin dielectric layer using ALD.
A decade of mathematical certainty: Reflections on the Automated Reasoning Group
In 2016 a small research team at Amazon announced our presence to the world with the launch of the Automated Reasoning Group (ARG). Our vision was bold: use mathematical logic to not just test AWS systems but to prove, with mathematical certainty, that they work correctly.