Array
(
    [content] => 
    [params] => Array
        (
            [0] => /forum/threads/is-nvidias-cuda-moat-cracking-this-could-be-amds-moment.25629/
        )

    [addOns] => Array
        (
            [DL6/MLTP] => 13
            [Hampel/TimeZoneDebug] => 1000070
            [SV/ChangePostDate] => 2010200
            [SemiWiki/EmailDomainReplace] => 1000010
            [SemiWiki/Newsletter] => 1000010
            [SemiWiki/WPMenu] => 1000010
            [SemiWiki/XPressExtend] => 1000010
            [ThemeHouse/XLink] => 1000970
            [ThemeHouse/XPress] => 1010570
            [XF] => 2031070
            [XFI] => 1060170
        )

    [wordpress] => /var/www/html
)

Is Nvidia's CUDA Moat Cracking? This Could Be AMD's Moment

karin623

Member
1785531152531.png


For more than a decade, AMD has been chasing Nvidia in AI—but CUDA has remained the industry's biggest barrier. That dynamic may finally be shifting.

This story explores why agentic AI could weaken CUDA's long-standing software moat, why AMD may be better positioned than ever to capitalize, and what this turning point could mean for the broader AI hardware ecosystem.

If CUDA's advantage is no longer untouchable, AMD's biggest opportunity may have finally arrived.

A transcript of a four-hour recording from DeepSeek founder Liang Wenfeng’s(梁文鋒) first fundraising meeting in May, held as the company prepared for a future IPO, has been leaked.

Bloomberg reported a few days ago that Liang was so angered by the leak that he suspended the second round of fundraising. To me, that effectively confirms the authenticity of the leaked conversation.

The remarkably candid private remarks from arguably the most important figure in China’s AI industry today contain a wealth of valuable information. Here are a few highlights.

 
Last edited by a moderator:
From my perspective, CUDA hasn't been the hardware moat for data center inference for at least the past year. With open models and open-sourced model serving engines like vLLM and SGlang, the real moat is the set of cost/power/throughput/latency Pareto curves for systems for a broad range of models. And the moat behind that is the speed and capability of each hardware supplier for system-level co-optimization for large MoE attention-based models. One great example here:


And a far better analysis of what AMD just rolled out at their AI conference and where they still might be challenged, here;


ps: thought the most important insight from Liang's transcript was the 4x lower performance plus "2 years later" penalty of Chinese / Huawei data center systems.
 
Last edited:
Back
Top