▾ G11 Media Network: | ChannelCity | ImpresaCity | SecurityOpenLab | Italian Channel Awards | Italian Project Awards | Italian Security Awards | ...
InnovationOpenLab

Cerebras Launches World’s Fastest Inference for Meta Llama 4

#AI--Cerebras Systems, the pioneer in accelerating generative AI, today announced the launch of Llama 4 inference across its CS-3 systems and on-demand cloud platform. Cerebras achieves over 2,600 tok...

Business Wire

Over 2,600 tokens per second performance enables instant AI interactions, real-time reasoning, blazing-fast code generation, and the next generation of agentic AI applications

SUNNYVALE, Calif.: #AI--Cerebras Systems, the pioneer in accelerating generative AI, today announced the launch of Llama 4 inference across its CS-3 systems and on-demand cloud platform. Cerebras achieves over 2,600 tokens per second on Llama 4 Scout – 19x faster than the fastest GPU solutions as verified by Artificial Analysis, a third-party AI benchmarking service. Llama 4 is available today on Cerebras via API access powered by our U.S.-based data centers, and via Hugging Face.

“Artificial Analysis has independently benchmarked Cerebras as achieving record speeds of over 2,600 output tokens per second on Llama 4 Scout, the highest we have measured across all providers. Cerebras also delivered the fastest end-to-end response time at just 0.5 seconds in our benchmarks,” said George Cameron, Co-Founder and Chief Product Officer at Artificial Analysis. “These speeds support AI developers to build agentic multi-step applications and real-time experiences that might be significantly limited on slower systems.”

“Llama 4 on Cerebras is now available on Hugging Face, giving developers access to faster inference for Meta’s latest open models,” said Clement Delangue, CEO of Hugging Face. “We’re thrilled to expand our partnership and bring this capability to millions of developers building the future of open AI.”

World’s Fastest Llama 4 Inference

Cerebras is widely recognized as the world’s fastest AI inference provider, achieving a record 2,500 tokens per second in Llama 3.3 70B. With this launch, Cerebras extends its performance leadership with the new Llama 4 family.

“Cerebras delivers the fastest Llama 4 Scout performance in the world,” said Andrew Feldman, CEO and co-founder of Cerebras. “We’ve broken all records delivering 2,600 tokens per second while the leading GPU solution delivers 137 tokens per second. Being 19 times faster than the competition opens up AI applications in real-time reasoning, agentic workflows that GPU-based platforms just can’t match.”

Instant AI, Instant Reasoning

Cerebras inference architecture stores all model parameters entirely in on-chip SRAM, delivering memory bandwidth far beyond traditional systems. This eliminates memory transfer bottlenecks and enables ultra-fast response times. The result is a breakthrough experience: instant AI, instant reasoning, and fast code generation—all at full model fidelity.

Llama 4

Llama 4 is Meta’s most advanced open-weight model family yet, introducing a Mixture of Experts (MoE) architecture that activates only a subset of parameters per token—dramatically improving efficiency while enhancing performance. The lineup includes Llama 4 Scout, with 17B active parameters and long-context capabilities, delivering state-of-the-art performance for its size, and Llama 4 Maverick, a 400B-parameter model with 128 experts that excels at multilingual understanding, creative writing, and general assistant tasks. These models offer developers powerful tools for building sophisticated, high-quality AI applications.

“Llama 4 is the most powerful Llama model to date, with better accuracy, improved multilingual capabilities, enhanced safety, and more,” said Karl Freund, founder and principal analyst at Cambrian AI. “By delivering over 2,600 tokens per second for Scout – more than 38 times faster than closed models from OpenAI and Anthropic – Cerebras is retaining the inference performance crown as the world’s fastest Llama provider. That means Meta and Cerebras are helping developers everywhere to move faster, go deeper, and build better than ever before, including personalized multimodal experiences.”

Availability

Meta has released two models in the Llama 4 family so far—the 109-billion parameter Llama 4 Scout and the 400-billion parameter Llama 4 Maverick. The Llama 4 Scout model is available immediately on Cerebras Cloud, Cerebras CS-3 hardware, and via Hugging Face. The Llama 4 Maverick model is coming soon.

About Cerebras Systems

Cerebras Systems is a team of pioneering computer architects, computer scientists, deep learning researchers, and engineers of all types. We have come together to accelerate generative AI by building from the ground up a new class of AI supercomputer. Our flagship product, the CS-3 system, is powered by the world’s largest and fastest commercially available AI processor, our Wafer-Scale Engine-3. CS-3s are quickly and easily clustered together to make the largest AI supercomputers in the world, and make placing models on the supercomputers dead simple by avoiding the complexity of distributed computing. Cerebras Inference delivers breakthrough inference speeds, empowering customers to create cutting-edge AI applications. Leading corporations, research institutions, and governments use Cerebras solutions for the development of pathbreaking proprietary models, and to train open-source models with millions of downloads. Cerebras solutions are available through the Cerebras Cloud and on-premises. For further information, visit cerebras.ai or follow us on LinkedIn or X.

Fonte: Business Wire

If you liked this article and want to stay up to date with news from InnovationOpenLab.com subscribe to ours Free newsletter.

Related news

Last News

RSA at Cybertech Europe 2024

Alaa Abdul Nabi, Vice President, Sales International at RSA presents the innovations the vendor brings to Cybertech as part of a passwordless vision for…

Italian Security Awards 2024: G11 Media honours the best of Italian cybersecurity

G11 Media's SecurityOpenLab magazine rewards excellence in cybersecurity: the best vendors based on user votes

How Austria is making its AI ecosystem grow

Always keeping an European perspective, Austria has developed a thriving AI ecosystem that now can attract talents and companies from other countries

Sparkle and Telsy test Quantum Key Distribution in practice

Successfully completing a Proof of Concept implementation in Athens, the two Italian companies prove that QKD can be easily implemented also in pre-existing…

Most read

Morning Walk Finds Its Stride as Mid-Market Brands Embrace Performance…

Morning Walk, the modern performance branding company known for building brands while driving scalable business results, announced today a surge in growth…

Dassault Systèmes and Airbus Extend Strategic Partnership to Use Virtual…

#3DEXPERIENCE--Dassault Systèmes (Euronext Paris: FR0014003TT8, DSY.PA) and Airbus have extended their long-term strategic partnership, putting the 3DEXPERIENCE…

CORRECTING and REPLACING Saviynt Hires Identity Veteran Roger Hsu to Accelerate…

Headline of release issued April 22, 2025, at 9:30 p.m. PT/April 23, 2025, at 12:30 a.m. ET should read: Saviynt Hires Identity Veteran Roger Hsu to Accelerate…

AppViewX Expert to Present Session on Achieving Crypto-Agility for the…

#CLM--AppViewX, a leader in automated certificate lifecycle management (CLM) and public key infrastructure (PKI) solutions, will take the stage at BSidesSF…

Newsletter signup

Join our mailing list to get weekly updates delivered to your inbox.

Sign me up!