Array
(
    [content] => 
    [params] => Array
        (
            [0] => /forum/threads/ians-take-on-the-new-cerebras-wafer-scale-engine-3-turbo-and-cs-4-rack.25733/
        )

    [addOns] => Array
        (
            [DL6/MLTP] => 13
            [Hampel/TimeZoneDebug] => 1000070
            [SV/ChangePostDate] => 2010200
            [SemiWiki/EmailDomainReplace] => 1000010
            [SemiWiki/HtmlMailer] => 1000200
            [SemiWiki/Newsletter] => 1000010
            [SemiWiki/WPMenu] => 1000010
            [SemiWiki/XPressExtend] => 1000010
            [ThemeHouse/XLink] => 1000970
            [ThemeHouse/XPress] => 1010570
            [XF] => 2031270
            [XFI] => 1060170
        )

    [wordpress] => /var/www/html
)

Ian's take on the new Cerebras Wafer Scale Engine 3-Turbo and CS-4 rack.

Daniel Nenni

Founder
Staff member
1787162869908.png


Cerebras has a new wafer scale chip and a new box to put it in. WSE 3.5 doubles the frequency and roughly doubles performance over WSE 3 on the same TSMC N5 node, and CS4 drops three of them into one 120 kilowatt rack using a backpack design for power, cooling and modular networking. I got an early look at media day, so here are the numbers, the rack teardown, and what it signals about the OpenAI agreement.

 
Definitely worth a listen...

Here is the AI version:

Cerebras’ Wafer Scale Engine 3 Turbo (WSE-3T) is an accelerated wafer-sized AI processor for high-speed inference and training. Fabricated as a single 46,225-square-millimeter device, it integrates four trillion transistors, 900,000 AI-optimized cores, and 44 GB of on-wafer SRAM. Each engine delivers 250 PFLOPS of AI compute, 43.2 petabytes per second of memory bandwidth, 53.5 petabytes per second of fabric bandwidth, and 2.4 terabits per second of external I/O. Compared with WSE-3, compute and bandwidth double, while I/O latency falls from five microseconds to as little as two.

The CS-4 combines three WSE-3T processors in Cerebras’ new Nexus rack-scale architecture. Together, they provide 750 PFLOPS, 129.6 petabytes per second of memory bandwidth, 160.5 petabytes per second of fabric bandwidth, and 7.2 terabits per second of I/O. Modular “Wafer-Scale Backpacks” integrate each wafer with power conversion, direct liquid cooling, controls, and networking, simplifying deployment and upgrades. Cerebras claims CS-4 delivers twice the speed and ten times the throughput per watt of CS-3, plus inference up to thirty times faster than GPU systems. Low-latency wafer links support interactive inference for models exceeding ten trillion parameters and clusters handling more than fifty trillion parameters. This makes the platform aimed squarely at hyperscale AI operators.

Sources: Cerebras CS-4 announcement, CS-4 product page
 
It's not a new chip. It's just the old chip overclocked

I don't think I'd call this overclocked:

- They changed the power delivery system to better feed the system
- The cooling is significantly upgraded
- The core is the same, but power consumption is only up 100-120%, while clock speed is up 100% (2.8 vs 1.4 GHz).

That last part indicates that 2.8 GHz is well within standard voltage spec for the chip/cores, it was just power and cooling limited from reaching higher clocks.
 
Back
Top