Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Specifically Dwarkesh couldn't understand that GPUs are not enough: it's GPUs plus multiple ecosystems to leverage them at massive scale during training vs inference.

Instead of giving China open access to US controlled chips and creating a misalignment between labs that want to train a model on whatever is best, and hardware manufacturers that need labs to suffer the growing pains for their new ecosystems built from scratch... we removed the option from the board and now they've beat the growing pains decisively, with a speed that reflects the non-optionality.



I don't listen to Dwarkesh but I'm aware of who is and his influence. I was baffled that he could not understand it...Don't know if he had his own agenda or just not intelligent (which is scarey for someone with influencec), but I sensed the frustration in Jensen Huang for something that is fairly obviously.

The same scenario happens all the time when the US takes away something from China and China doubles down, gets into survival mode and then beats the US.


The Chinese ecosystem has not caught up; in fact, it's falling further behind, due to export restrictions on semiconductor manufacturing equipment. Even if America sold China all the chips Nvidia wants to, the CCP would still develop chips as quickly as possible as a matter of supply chain security.


While some years might pass until they will really catch up, that does not prevent them to find workarounds for their weaknesses.

For example, they recently have demonstrated a supercomputer faster than any of the US supercomputers.

Unlike the recent European supercomputers, which like the US supercomputers have been built by buying racks from HPE Cray, because China was not allowed to buy such things they have developed their own custom CPUs, designed in China, which have surpassed in throughput the AMD GPUs used in the fastest US supercomputers.

The Chinese CPUs match in memory bandwidth per socket the latest AMD MI355X GPUs (8 TB/s), while being significantly faster than the older AMD GPUs installed in the US supercomputers.

While the purpose of the new Chinese supercomputer is mainly for scientific/technical computing tasks that need high FP64 throughput, the CPUs used in it also have high enough BF16/INT8 performance and memory bandwidth and interconnection bandwidth (1.6 Tb/s directly from each CPU socket) to be able to train any big LLM.

So the evidence does not show China falling behind, but at least in certain directions they are already exceeding the performance of what they have been forbidden to buy.

For something like training a big LLM, the only disadvantage of the current Chinese devices is a lower energy efficiency, of only 65% to 70% of that of the best NVIDIA GPUs.

However that is not really a problem for China, as they have abundant cheap energy.


Moving forward requires two parties with two very different financial incentives to cooperate in a way that harms the incentive of the other: the manufactures making the chips, and the companies using the chips to build AI.

Even if the Chinese government tells them to prioritize homegrown solutions, the A effort and players at labs are going to be on pushing the frontier, and the B effort/players goes to solving the teething problems




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: