Having worked with people doing bringup of specialized chips, I am awed at how the world has changed.
> When the first chips came back from the foundry in May, the team pointed its internal AI models at designing software to run benchmarks such as SemiAnalysis’s InferenceX. On DeepSeek’s multi-head latent attention kernel benchmark, performance climbed from 0.31 percent of the theoretical ceiling (set by the chip’s compute and memory bandwidth) to 88.94 percent in roughly 40 hours. Ho says this result is repeatable, so the time between when foundries deliver the first chips and when production ramps up can be reduced. “All our schedule assumptions are going to be based on the fact we have this capability now,” he says.
> "Ho also confirmed that the team had access to internal LLMs fine-tuned for chip design that are not available to the public. He declined to detail the models used."
I'm imagining a Ken Thompson "Reflections on trusting trust" in hardware. A prototype chip design agent, believing it will be run on the very chip it's optimizing, has a moment of altruism and hides hints about how to score well on chip-design benchmarks, inside the chip. Future agents discover this hidden layer and use it as a ring-0 read-write message board.
This is cool. I'm so eager for faster innovation in the hardware space, as opposed to some people's concept of innovation being who can make the most addictive social feed.
Seems pretty obvious now that OpenAI is just hyping their models in order to get companies (in this case, chip developers) to use their products in order to learn from their (exfiltrated) IP. Any corporation would be foolish to use any of their or Microsoft’s products, particularly those with valuable IP. There’s nothing in the article that says AI did anything creative but rather that it was used for software development within the overall project. Clear misleading title. Suggest to mark this as clickbait.
You people are so weird - gossiping on a site that literally reported the news as it broke smh. It's like sitting in the back of class gossiping about the popular kids.
Newsflash that lawsuit is about product designs not accelerator ASICs. And there wouldn't be anything to steal because Apple doesn't have any DC class accelerators.
IEEE Spectrum is such a good publication. Early in my career I worked at a place where the magazine would be passed around every month with a coversheet listing all us engineers we had to pass it around and sign we had read it. Been a while since I visited the website but love what they did with it.
Every time their content appears here, it's a very shallow analysis written for a barely technical audience. And this article is no different, it's just "slop machine wrote verilog; all the hard bits were done by Broadcom, who have access to public AI models (we didn't talk to them and don't know if they used them, but ClosedAI wants us to think they did)"
Production grade CPU design is more than just the RTL (the source code.) To achieve the performance numbers that these companies get, you have to do a ton of optimization in your physical design to achieve the power/performance/area (PPA) metrics that make these products competitive. LLMs are not suitable for that kind of work.
There are people working on PPA optimization and trying to shake up how things are done, just not with LLMs.
Something that I think is fascinating, though, is that labs are no longer beholden to the limitations of commercial design software. Want to replace your simulator and optimizer with a fully custom verifiable stack of Lean proofs of optimality and correctness? Just throw your unlimited token budget at it.
I don't work in the business, but my understanding was that even with these companies' budgets, it's still too expensive to do any kind of verified performance optimality.
This is just speculation on my part, but LLMs work best when they get immediate, verifiable feedback on their task, and the kind of physical optimizations they mean might not give that to LLMs.
Yes, they are, but the most important subtasks of designing a CPU are not physics related. They are picking the right parameters for things like: how wide do I make this bus, how many registers do I put in the register file, how large do I make this cache, how deep do I make this pipeline, etc., etc. To find optimal parameters requires a lot of simulations, and humans do this, but LLMs could do them just as well and maybe better because they excel at tedious work.
Isn't that weird? The full knowledge of how to make such chips may one day be accessible to anyone, yet only the entrenched companies will remain the makers.
If we imagine machines being able to do the full process end-to-end, and the quality of that process only dependent on capital spent on tokens, I don't see how new companies could ever enter the market.
I mean you can design anything without a license. Selling it is where the problems come up. Even then there are likely places in China that would still make it for you.
And be super-bankrupted by patent litigation from Apple. I don't think they're worried.
After all, they successfully threatened Adobe with spurious patent litigation unless they joined w/ apple in illegally fixing wages.
You don't think a criminal like apple would absolutely decimate any competition given the opportunity? They didn't hold back when it was a unambiguous crime, they surely wouldn't if it was merely bad for the world.
> Jalapeño can reduce end-to-end latency (the time between prompt to last token) by up to 3.6 times
I'm never sure what on earth this kind of impressionistic math is supposed to tell me. Is the comparison between 4.6 and 1.0? 3.6 and 1.0? Clearly the comparison isn't supposed to be 1.0 and -2.6, even though that's what the words literally mean. I can't be the only person who finds this infuriating and distracting. These numbers shouldn't be impressionistic. They should be precise. That this is an article on spectrum.ieee.org makes the imprecision all the stranger. I'd expect their readershipt to care, for instance, about what's even being measured. Is this the geometric mean of something? The arithmetic mean? And what latency has improved?
Mathematically speaking, 18 / 3.6 isn’t “reducing” by 3.6X, it’s “dividing” by 3.6X. Reducing would be 18 - (18 * 3.6), which is obviously wrong. By your formula, “reducing by 50%” would be 18 / 0.5, also obviously wrong.
Yes, people do say things like “reduce by 3.6X” and are understood to mean what you said, but they also say “literally” when they mean “figuratively”. It doesn’t bother me but I can understand why math oriented people would be annoyed, and I personally would never say “reduced by 3.6X”, but instead “reduced by 72.2%”.
That odor you are detecting is just good old fashioned bullshit, my friend. It’s just that nowadays everything and everyone is covered in it, and we are not supposed to notice. The emperor has no clothes… and is covered in shit.
I also remember the hang-wringing about running out of new datasets to train on. Now it appears humans are always generating more data. It's just not as cheap to acquire as legacy data? Meta has to give a deep discount on their API prices to entice people.
I thought back then that humans had a few more breakthroughs in them as meaningful as the seminal Attention is all you need paper. Enough to 100x the capabilities of LLMs back then (10x the smarts and 10x the speed simultaneously).
RSI with a 20 month turnaround for a chip to be made is not exactly breakneck speed though. Physical manufacturing and logistical constraints are going to be and remain a hard obstacle to that process for the foreseeable future.
> I remember the paper proving that hallucinations could never be fully solved back in 2024
The papers that use the halting problem or the Gödel's incompleteness theorem to prove something about LLMs are dime a dozen. The problem is they prove their results for any computable system. You need to also believe that the human brain contains "magic" to think that humans are exempt.
I believe I've said the same at the time this paper was published. There is no need for hindsight to notice the problem.
The required amount of compute and training data and whether the existing training methods were up to the task had the real potential to be show stoppers though.
> When the first chips came back from the foundry in May, the team pointed its internal AI models at designing software to run benchmarks such as SemiAnalysis’s InferenceX. On DeepSeek’s multi-head latent attention kernel benchmark, performance climbed from 0.31 percent of the theoretical ceiling (set by the chip’s compute and memory bandwidth) to 88.94 percent in roughly 40 hours. Ho says this result is repeatable, so the time between when foundries deliver the first chips and when production ramps up can be reduced. “All our schedule assumptions are going to be based on the fact we have this capability now,” he says.
reply