▾ G11 Media Network: | ChannelCity | ImpresaCity | SecurityOpenLab | Italian Channel Awards | Italian Project Awards | Italian Security Awards | ...
InnovationOpenLab

Meta Collaborates with Cerebras to Drive Fast Inference for Developers in New Llama API

Meta has teamed up with Cerebras to offer ultra-fast inference in its new Llama API, bringing together the world’s most popular open-source models, Llama, with the world’s fastest inference techno...

Business Wire

SUNNYVALE, Calif.: Meta has teamed up with Cerebras to offer ultra-fast inference in its new Llama API, bringing together the world’s most popular open-source models, Llama, with the world’s fastest inference technology, delivered by Cerebras. This new platform unlocks groundbreaking possibilities for a massive developer audience.

Developers building on the Llama 4 Cerebras model in the API can expect generation speeds up to 18 times faster than traditional GPU-based solutions. This acceleration unlocks an entirely new generation of applications that are impossible to build on other technology. Real-time agents, conversational low latency voice, interactive code generation, and instant multi-step reasoning — all of which require chaining multiple LLM calls — can now be completed in seconds rather than minutes.

By partnering with Meta to serve Llama models from Meta’s new API service, Cerebras gains exposure to an expanded global developer audience and deepens its business and partnership with Meta and their incredible teams.

Since launching its inference solutions in 2024, Cerebras has delivered the world’s fastest Llama inference, serving billions of tokens through its own AI infrastructure. The broad developer community now has direct access to a robust, OpenAI-class alternative for building intelligent, real-time systems — backed by Cerebras speed and scale.

“Cerebras is proud to make Llama API the fastest inference API in the world,” said Andrew Feldman, CEO and co-founder of Cerebras. “Developers building agentic and real-time apps need speed. With Cerebras on Llama API, they can build AI systems that are fundamentally out of reach for leading GPU-based inference clouds.”

Cerebras is the fastest AI inference solution as measured by third party benchmarking site Artificial Analysis, reaching over 2,600 tokens/sec for Llama 4 Scout compared to ChatGPT at ~130 tokens/sec and DeepSeek at ~25 tokens/sec.

Developers will be able to access the fastest Llama 4 inference by selecting Cerebras from the model options within the Llama API. This streamlined experience will make it easy to prototype, build, and scale real-time AI applications. To sign up for early access to the Llama API and to experience Cerebras speed today, visit www.cerebras.ai.

About Cerebras Systems

Cerebras Systems is a team of pioneering computer architects, computer scientists, deep learning researchers, and engineers of all types. We have come together to accelerate generative AI by building from the ground up a new class of AI supercomputer. Our flagship product, the CS-3 system, is powered by the world’s largest and fastest commercially available AI processor, our Wafer-Scale Engine-3. CS-3s are quickly and easily clustered together to make the largest AI supercomputers in the world, and make placing models on the supercomputers dead simple by avoiding the complexity of distributed computing. Cerebras Inference delivers breakthrough inference speeds, empowering customers to create cutting-edge AI applications. Leading corporations, research institutions, and governments use Cerebras solutions for the development of pathbreaking proprietary models, and to train open-source models with millions of downloads. Cerebras solutions are available through the Cerebras Cloud and on-premises. For further information, visit cerebras.ai or follow us on LinkedIn, X and/or Threads.

Fonte: Business Wire

If you liked this article and want to stay up to date with news from InnovationOpenLab.com subscribe to ours Free newsletter.

Related news

Last News

RSA at Cybertech Europe 2024

Alaa Abdul Nabi, Vice President, Sales International at RSA presents the innovations the vendor brings to Cybertech as part of a passwordless vision for…

Italian Security Awards 2024: G11 Media honours the best of Italian cybersecurity

G11 Media's SecurityOpenLab magazine rewards excellence in cybersecurity: the best vendors based on user votes

How Austria is making its AI ecosystem grow

Always keeping an European perspective, Austria has developed a thriving AI ecosystem that now can attract talents and companies from other countries

Sparkle and Telsy test Quantum Key Distribution in practice

Successfully completing a Proof of Concept implementation in Athens, the two Italian companies prove that QKD can be easily implemented also in pre-existing…

Most read

U.S. Data Center Construction Market Outlook Report 2025-2030 Featuring…

The "U.S. Data Center Construction Market - Industry Outlook & Forecast 2025-2030" report has been added to ResearchAndMarkets.com's offering. The…

Alibaba Group Announces March Quarter 2025 and Fiscal Year 2025 Results

$BABA #alibaba--Alibaba Group Holding Limited (NYSE: BABA and HKEX: 9988 (HKD Counter) and 89988 (RMB Counter), “Alibaba”, “Alibaba Group” or the “company”)…

Australia Social Commerce Intelligence Databook 2025: An $8.58 Billion…

The "Australia Social Commerce Market Intelligence and Future Growth Dynamics Databook - 50+ KPIs on Social Commerce Trends by End-Use Sectors, Operational…

J.D. Power Names Joshua Peirez New CEO

J.D. Power today announced that Joshua Peirez will assume the role of President and CEO of J.D. Power, guiding the company in its next phase of growth…

Newsletter signup

Join our mailing list to get weekly updates delivered to your inbox.

Sign me up!