MLCommons
Online/Digital
MLCommons is an Artificial Intelligence engineering consortium, built on a philosophy of open collaboration to improve AI systems. Through our collective engineering efforts with industry and academia we continually measure and improve the accuracy, safety, speed, and efficiency of AI technologies–helping companies and universities around the world build better AI systems that will benefit society. Source
Actions
Media Outlet details
| Scope | International, Trade/B2B |
|---|---|
| Language | English |
| Country | United States of America |
|
Similarweb UVM |
Request pricing |
|
Comscore UVM |
Request pricing |
Recent Articles
Search ArticlesMLPerf Training Introduces Its First LLM Post-Training Benchmark
MLPerf Training is adding a new LLM Post-Training benchmark beginning with the v6.1 submission round in October 2026, complementing the existing suite of pre-training benchmarks. While pre-training builds foundational intelligence via massive-scale data ingestion, post-training is a complementary process that refines that foundation to excel at specific tasks. Since the second half of 2025, significant advances in LLM training have come from scaling post-training.
MLCommons Joins EU-Funded AIRIS Project to Build and Benchmark Next-Generation Biomedical AI
Over the summer, MLCommons was honored to be announced as a consortium member in a new research project AIRIS, (Mechanism-Informed Multimodal Generative AI for Causal and Dynamical Modeling in Biomedical Research), funded by the European Union’s Horizon Europe Program. AIRIS unites 21 partners from Europe, Canada and the US to develop generative AI models that integrate biological knowledge and clinical data.
Where the Industry Is Investing: A Look at MLPerf Inference v6.1
Every round of MLPerf® Inference is a snapshot of where the industry is investing its engineering energy, and v6.1 stands out on two fronts: it is the broadest field of submitters we have ever seen, and it marks a clear inflection toward agentic and end-to-end benchmarking alongside a wave of newly submitted hardware. The highlights of this round include a record number of submitters, new accelerators that significantly improve per-device performance, and a new multi-turn benchmark.
MLCommons Sets Participation Record with New MLPerf Inference v6.1 Benchmark Results
MLCommons® announced new results for its industry-standard MLPerf® Inference v6.1 benchmark suite. This release, which set a new high-water mark for the number of submitting organizations, introduces two new tests aligned with recent AI inference deployment trends. It also features the first peer-reviewed performance results for several recently released or soon-to-be-released AI platforms, demonstrating up to a 5.7X performance gain compared to just one year ago.
MLCommons Agent Reliability Profile Named a Finalist in Global Agentic Regulator Hackathon
How do you know whether an AI agent will stay in its lane? It’s a question financial institutions and their regulators can’t yet answer with confidence — and it’s the question at the center of the MLCommons Financial Services Working Group’s Agent Reliability Profile, which has been selected as a finalist in the C:>DIR Global “Agentic Regulator” Hackathon. AI agents don’t just answer questions — they take actions.
MLCommons Releases New MLPerf Storage v3.0 Benchmark Results
Today, MLCommons® announced the results of its industry-standard MLPerf® Storage v3.0 benchmark suite, which measures the performance of storage systems for machine learning (ML) workloads in an architecture-neutral, representative, and reproducible manner. Version 3.0 expands the tests in the suite to represent the breadth of storage workloads that AI systems can generate, and adds support for an S3 object storage access layer alongside the existing POSIX layer.
The key to trustworthy AI evaluation is secrecy by design
Every business that deploys AI will want to know: does this system perform reliably and safely for my use case and with my data? But it’s not as simple as hooking everything together and running a test. Deployers, like banks, need to keep sensitive data safe. AI solution providers, like frontier labs, are very protective of model weights.
Introducing the MLPerf End-to-End RAG Inference Benchmark
The MLCommons MLPerf Inference Working Group is excited to introduce the first instance of a new End-to-End Retrieval-Augmented Generation (RAG) benchmark. By answering from documents retrieved at query time rather than from weights alone, RAG reduces hallucination and draws on current, private knowledge, which has made it one of the most common ways language models are deployed.
MLPerf Client v2.0 Expands AI PC Benchmarking with Image Generation and Agentic AI
MLCommons®, an open engineering consortium dedicated to improving machine learning performance and transparency, today announced the release of MLPerf® Client v2.0, the latest version of its industry-standard benchmark for evaluating AI performance on personal computers. MLPerf Client measures how effectively PCs—from laptops and desktops to workstations—run AI workloads locally.
How to Tell When a Benchmark Is Worth Trusting
If you work in enterprise AI, you’ve probably been here: a vendor claims their model leads on a popular leaderboard. A procurement team is weighing two systems and benchmark scores are the tiebreaker. An executive deck cites a safety benchmark to argue a model is “production-ready.” Some of those numbers are real. Some are what we’d call benchmark washing – the selective use of convenient results to imply performance, reliability, safety, or readiness that the evidence doesn’t actually support.