• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Announcement

Cerebras is working on wafer-scale stacked DRAM for larger AI models

A Cerebras veteran says the company chose SRAM for its bandwidth. Its stacked DRAM work aims to add memory capacity while preserving fast inference.

2 Sources, 10d ago, first seen 10d ago

TLDR

In part two of an AMA, a Cerebras veteran says the company is working on wafer-scale stacked DRAM to give larger AI models more capacity while preserving fast inference. The work also raises questions about how to power, cool and reliably manufacture it.

Combined views

—

2 Sources, first seen 10d ago

— likes— comments— saves— reposts

Combined views

—

2 Sources, first seen 10d ago

— likes— comments— saves— reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Related

OpenAI's GPT6.1 Sol Ultrafast reportedly runs on Nvidia GPUs, not Cerebras

SemiAnalysis claims the model is running at a low batch size on Nvidia GPUs and asks whether Cerebras might serve it later.

Cerebras-built AI assistant claimed to be 19x faster than three others on a dinner-reservation task

Cerebras says its assistant used Qwen 3.8 27B running at about 1,500 tokens per second. It compared the assistant with Grok Bot, Meta Muse and Claude Cowork on the same dinner-reservation task.

An unnamed buyer reportedly secured $400M in debt financing and is buying Cerebras machines to run alongside Nvidia GPUs

A user says Cerebras puts memory directly on its chips, speeding data movement during token generation and, in turn, inference.

2 Sources

Sean Lie@seanlieI joined AMD right after graduating from MIT. Two stints there and a decade building @cerebras later, I still lose sleep over how much memory and compute to put on a chip. I’ve spent my entire career in computer architecture, and I’ve never seen the problems change this quickly. Models keep getting larger. The demands on the hardware keep growing. Decisions about where to put memory and how to connect it have become central to what people can build with AI. We chose SRAM for its bandwidth. Now we’re working on wafer-scale stacked DRAM to give larger models more capacity while preserving fast inference. That brings us back to some very physical problems: how to power it, cool it, and manufacture it reliably. This is the work I love. In part two of our AMA, I answered your questions about those decisions and where we’re going next. 0:30 What changes between CS4, CS5, and CS6? 2:19 HBM already stacks memory. What’s different here? 3:34 We chose SRAM for speed. Why add DRAM? 4:52 Drawing the stack is easy. Making it work isn’t. 7:23 The memory-versus-compute decision that keeps me up at night10d
Sarah Chieng@MilksandMatchaRT @seanlie: I joined AMD right after graduating from MIT. Two stints there and a decade building @cerebras later, I still lose sleep over…10d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    Cerebras

    2 Sources

    Sean Lie@seanlieI joined AMD right after graduating from MIT. Two stints there and a decade building @cerebras later, I still lose sleep over how much memory and compute to put on a chip. I’ve spent my entire career in computer architecture, and I’ve never seen the problems change this quickly. Models keep getting larger. The demands on the hardware keep growing. Decisions about where to put memory and how to connect it have become central to what people can build with AI. We chose SRAM for its bandwidth. Now we’re working on wafer-scale stacked DRAM to give larger models more capacity while preserving fast inference. That brings us back to some very physical problems: how to power it, cool it, and manufacture it reliably. This is the work I love. In part two of our AMA, I answered your questions about those decisions and where we’re going next. 0:30 What changes between CS4, CS5, and CS6? 2:19 HBM already stacks memory. What’s different here? 3:34 We chose SRAM for speed. Why add DRAM? 4:52 Drawing the stack is easy. Making it work isn’t. 7:23 The memory-versus-compute decision that keeps me up at night10d
    Sarah Chieng@MilksandMatchaRT @seanlie: I joined AMD right after graduating from MIT. Two stints there and a decade building @cerebras later, I still lose sleep over…10d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet