AMD has struck a deal to buy Taalas, a Toronto startup whose engineers etch a neural network's weights straight into a silicon chip. The acquisition was announced on August 6; neither side disclosed a price, citing only “customary closing conditions” and pending regulatory approval.
A standard GPU spends every inference cycle hauling billions of parameters from memory to its compute cores, and that shuttle is the real bottleneck when running large language models. Taalas skips it entirely: the weights aren't stored separately, they're burned into the transistors themselves. Its demo chip, the HC1, is an 815-square-millimeter die packing 53 billion transistors on TSMC's 6-nanometer process. By the company's own benchmarks it hits up to 17,000 tokens per second per user running Llama 3.1 8B, beating Nvidia's H200 and B200 plus chips from Groq, SambaNova and Cerebras in that test.
The tradeoff is flexibility. A rack built around one model can't simply be repointed at another — the chip effectively is the model. Taalas says updating the weights only requires swapping two photomask layers rather than redesigning the whole chip, which cuts iteration costs, but every change still means a new fabrication run. That math only pencils out for models that run at massive, stable volume rather than ones that get replaced every couple of months.
AMD frames the purchase as filling out its inference lineup. “We're building a full-stack AI platform that gives customers the flexibility to deploy the right compute for every workload,” said Vamsi Boppana, AMD's senior vice president for its AI group. The company plans to fold Taalas' technology into its accelerator roadmap alongside Instinct GPUs, EPYC processors, the Helios rack platform and ROCm software — the same lineup AMD used in July to challenge Nvidia head-on. Taalas co-founder and CEO Ljubisa Bajic, whose Toronto-based company was founded in 2023, called the deal a way to speed up development while keeping the team's Canadian roots.
For now this is a demo chip, not a shipping product, and the price tag stays private. But betting on model-specific silicon runs against the market's own rhythm, where new LLM versions land every few weeks — AMD is wagering that cheap, dedicated inference makes sense precisely where models have already stopped changing so fast.



