MLCommons
Online/Digital
MLCommons is an Artificial Intelligence engineering consortium, built on a philosophy of open collaboration to improve AI systems. Through our collective engineering efforts with industry and academia we continually measure and improve the accuracy, safety, speed, and efficiency of AI technologies–helping companies and universities around the world build better AI systems that will benefit society. Source
Actions
Media Outlet details
| Scope | International, Trade/B2B |
|---|---|
| Language | English |
| Country | United States of America |
|
Similarweb UVM |
Request pricing |
|
Comscore UVM |
Request pricing |
Recent Articles
Search ArticlesMLPerf Endpoints v0.7: A Foundation Release
MLPerf has been the standard-bearer for measuring AI system performance since its launch in 2018. During that time, MLPerf has tracked over a 100X improvement in inference performance per watt for large language models and over a 50X improvement in training speed [1] [2]. Over the past eight years, the AI industry has matured, with AI services now used daily by enterprises and consumers worldwide.
MedPerf Meets Google Cloud Confidential Computing: Secure AI Benchmarking for Brain Tumor Research
During Google Cloud Next 2026 in Las Vegas, the MLCommons Medical AI working group and Google Cloud announced the enablement of MedPerf, MLCommons’ federated benchmarking orchestrator, on Google Cloud’s confidential compute capabilities. Both teams demonstrated this integration on a compelling real-world clinical use case: brain tumor segmentation. A brain tumor (glioblastoma) is a rare disease with devastating outcomes for life expectancy.
Call for Submission: Edge Agentic Inference Benchmark for MLPerf Inference v6.1
The MLCommons Edge LLM Taskforce is excited to introduce a new Edge Agentic Inference benchmark for the MLPerf Inference v6.1 round. As agentic LLMs – coding copilots, robotics controllers, and private on-prem assistants – increasingly run on-device, measuring how well and how fast these models call tools under a real edge budget is more important than ever.
Agentic Inference for MLPerf Inference
The MLPerf Inference benchmark suite must evolve alongside AI deployment patterns. Early inference benchmarks focused on image classification, object detection, speech recognition, recommendation, and single-turn language generation. Those workloads remain important, but they no longer cover one of the fastest-growing ways large language models are used in production: multi-turn agentic inference. For example, a coding assistant is far more complex than a single query.
The Benchmark Behind the Next Wave of Ultra-Low-Power AI
Machine learning (ML) is no longer confined to data centers and is transforming the world around us, adding more intelligence to our day-to-day lives. It now runs on doorbell cameras, hearing aids, factory sensors, and battery-powered wearables. These devices operate on a few milliwatts and must respond in real time. As that footprint expands, a hard question follows. How do you fairly measure their performance and efficiency when no two of these devices look alike?
MLCommons Releases MLPerf Client v1.6 with Performance Optimizations and Enhanced User Experience
MLCommons®, the open engineering consortium behind the industry-standard MLPerf® benchmarks, today announced the release of MLPerf Client v1.6, the latest update to its benchmark suite for evaluating AI performance on personal computers. MLPerf Client measures how effectively PCs—from laptops and desktops to workstations—run AI workloads such as large language models (LLMs) locally.
Chakra Comes of Age: A Standardized Trace Ecosystem for AI Systems Benchmarking and Co-design
When MLCommonsannounced the Chakra working group in July 2023, the premise was simple but ambitious: AI systems are moving too fast for the traditional benchmarking and co-design playbook. Production workloads live behind walls of proprietary code and models. Simulators, emulators, and replay tools each invent their own representations.
The patch model is breaking. AI evaluation needs a new way to disclose what it finds.
For about thirty years the security community has relied on a well-understood approach for handling dangerous findings. Coordinated vulnerability disclosure is a standard practice for a reason, and it can neatly solve hazard disclosure problems with transparency and technical rigor. A security researcher finds a flaw, reports it privately to the vendor, the vendor ships a fix, deployers update, and only once that window has closed do the details go public.
MLCommons Releases New MLPerf Inference v6.0 Benchmark Results
Today, MLCommons® announced new results for its industry-standard MLPerf® Inference v6.0 benchmark suite. This release includes several important advances that ensure the benchmark suite tests current, real-world scenarios for AI deployments and delivers a comprehensive picture of AI system performance. Five of the eleven datacenter tests in MLPerf Inference v6.0 are new or updated, and the release also includes a new object-detection test for edge systems.
A new GPT-OSS benchmark and DeepSeek R1 updates for latency-optimized reasoning
The MLPerf® Inference v6.0 release marks a significant expansion in our coverage of the open-weight large language model (LLM) landscape. As the industry moves toward more specialized and capable open models, the benchmarks must evolve to reflect these shifts in deployment strategies and model architectures.