
DeepSeek and Huawei’s TileLang toolkit targets CUDA lock-in
Published by AINave Editorial
DeepSeek and Huawei released a free, open-source six-module toolkit for Huawei Ascend 950 accelerators on September 30. Its most notable feature is TileLang, a language for writing AI kernels that can target Ascend as well as Nvidia, AMD and Apple hardware. The practical promise is less backend-specific kernel work, not a magic converter for existing CUDA code or evidence that the chips deliver equivalent performance.
TileLang changes the level at which developers program
CUDA asks developers to manage details such as threads, memory access and synchronization. TileLang instead uses tiles, or blocks of data, as its basic unit: programmers describe the computation, while the compiler handles scheduling and synchronization for the selected hardware backend, according to the account of TileLang’s design.
That distinction matters for portability. Developers who write kernels in TileLang can target Nvidia CUDA, AMD ROCm, Apple Metal and Huawei Ascend 950. The release is not evidence that arbitrary CUDA kernels can be reused unchanged. A team with a CUDA codebase would need to work in the TileLang programming model to benefit from its cross-backend approach.
Six modules cover different parts of the stack
TileLang is the programming layer; five accompanying modules address specific operations. DeepGEMM-Ascend handles matrix multiplication, while DeepEP-Ascend handles communication across devices in a cluster. TileKernels covers vector operations and memory access, FlashMLA focuses on long-context processing, and DeepSelect handles data filtering, according to the module descriptions.
The package mirrors tools DeepSeek had previously released for Nvidia hardware. That gives developers a set of components aimed at common AI workloads on Ascend, rather than a single language that by itself supplies every part of a production stack. DeepSeek says it uses TileLang as its primary tool for AGI research, but that statement does not establish how well every module performs on Ascend in other teams’ workloads.
A software route is not proof of a training replacement
The release addresses the cost of adapting kernels to different hardware. It does not demonstrate that Ascend can match Nvidia in demanding training workloads. The source says DeepSeek uses Huawei hardware for inference at scale but continues to use Nvidia for training, and that an earlier attempt to train on Huawei silicon ran into technical difficulties.S1
It also attributes to DeepSeek founder Liang Wenfeng an estimate that four Huawei GPUs equal one Nvidia GPU in effective compute, with Huawei about two years behind. That is an attributed comparison, not an independently audited benchmark, and it does not establish performance across workloads.S1
So the toolkit’s near-term significance is a lower software barrier to trying Ascend, not a demonstrated end to Nvidia dependence. Whether shared kernels and an open developer community can improve the experience over time will depend on implementation and performance across real workloads, not portability alone.






















