AIREITER

DeepSeek Ascend Infrastructure Components, Mapped to NVIDIA

Last Updated: 2026-09-30 19:10:22

A public GitHub repository can make a hardware stack look more available than it is. DeepSeek’s Ascend work is real and useful, but it is not a one-command replacement for an NVIDIA cluster: the current release centers on Ascend 950 kernels and needs CANN, torch_npu, and compatible hardware.

The short answer: open components, not a turnkey Ascend cluster

DeepSeek has published Ascend-specific code, including DeepGEMM-Ascend, an MIT-licensed kernel library that keeps the DeepGEMM API shape while targeting Huawei NPUs. The repository’s initial release supports Ascend 950 devices and documents CANN 9.20, torch_npu, Python 3.10+, a C++20 toolchain, and TileLang as part of the environment.

That is enough for a team with compatible Ascend access to inspect, build, benchmark, and integrate selected kernels. It is not a complete training or production stack.

Ascend-equipped labs, cloud operators, and enterprise teams can evaluate the code now. An NVIDIA-only developer can study it or use DeepSeek’s NVIDIA-oriented projects, but cannot run these Ascend kernels on an H100 or consumer GeForce card.

What DeepSeek actually released

DeepSeek’s open-infra-index groups its infrastructure work into layers. The original index is mostly NVIDIA/Hopper-oriented: FlashMLA is an MLA decoding kernel for Hopper GPUs, DeepEP is an expert-parallel communication library, DeepGEMM is an FP8 GEMM library, DualPipe and EPLB address distributed parallelism, and 3FS/Smallpond address data access.

The Ascend-specific release changes the hardware target for selected compute paths rather than replacing every layer at once. The clearest public artifact is DeepGEMM-Ascend, which covers:

  • BF16, FP8, and FP4 GEMM;
  • MQA logits;
  • grouped GEMM and MegaMoE paths;
  • an mHC prenorm kernel; and
  • Ascend-specific JIT compilation and layout transforms.

The project says it is API-compatible with DeepGEMM. That compatibility is valuable for model engineers, but it does not make CUDA binaries portable. Ascend’s matrix layouts, scaling-factor packing, compiler, runtime, and device-management APIs still matter.

Component map: Ascend equivalents for an NVIDIA-oriented stack

The table below maps roles, not claims of identical implementations. A counterpart performs a similar job; it may use a different API, communication fabric, or kernel strategy.

LayerConfirmed Ascend code or dependencyNVIDIA-oriented counterpartBoundary
Matrix multiplicationDeepGEMM-AscendDeepGEMM plus CUDA/Tensor Core kernelsSame GEMM role; different hardware primitives and data layouts
MoE grouped computeM-Grouped GEMM and MegaMoE in DeepGEMM-AscendDeepGEMM MoE layouts plus custom CUDA kernelsFused expert compute, with platform-specific shapes and limits
Expert dispatchNo complete DeepSeek Ascend dispatch library is identified in this release; use Ascend collectives and operator integrationsDeepEP with NVLink/RDMASimilar systems problem, not interchangeable repositories
Attention/logitsMQA logits in DeepGEMM-AscendFlashMLA for HopperSimilar workload; FlashMLA is explicitly Hopper-focused
PyTorch device bridgeTorchNPU (torch_npu)PyTorch CUDA backend, CUDA runtime, and cuBLASTorchNPU invokes Ascend NPUs; it does not emulate CUDA
Compiler/operator toolchainCANN and Ascend C/Bisheng toolingCUDA Toolkit, NVCC, PTX, cuBLAS, TritonPython code may look familiar while the device contract changes
Distributed runtimeTorchNPU collectives plus Huawei/operator deployment toolsNCCL, CUDA-aware networking, and NVIDIA cluster softwareFabric, drivers, collectives, and framework versions remain separate
Storage/data pathNo Ascend-specific DeepSeek storage component is identified here; storage is operator-providedDeepSeek’s 3FS/Smallpond plus an operator’s storage stackThe kernel release does not include a matching storage cluster

In short, each Ascend counterpart fills a systems role; it does not erase the NVIDIA software ecosystem.

Compute kernels: DeepGEMM-Ascend versus DeepGEMM

DeepGEMM-Ascend is the most concrete bridge in the release. Its README describes a lightweight abstraction over Ascend MAD primitives, hiding fractal layouts, alignment constraints, address calculations, and low-level parameters. It also uses Ascend-specific techniques such as sparse data loading and coroutine-based pipelining.

The published requirements are specific: Ascend 950-series hardware, CANN 9.20, torch_npu, Python 3.10 or newer, a C++20-compatible standard library, TileLang, and build dependencies including Tree-sitter. The documented install path includes git clone --recursive followed by pip install . --no-build-isolation.

The repository reports dense-GEMM utilization up to 99.8% of the stated hardware limit on an Ascend 950DT test setup. One BF16 case is listed at 431 TFLOPS against a 432-TFLOPS hardware limit; FP8 cases are listed at 861 against 865. These are kernel results for selected shapes, not end-to-end DeepSeek throughput or proof of parity with an NVIDIA cluster.

MoE execution: MegaMoE and expert dispatch

DeepSeek’s open-infra-index identifies expert-parallel infrastructure for its V3/R1 systems, so the important question is not only how fast one matrix multiply runs. Tokens must be routed, experts must compute, and results must be combined across ranks.

DeepGEMM-Ascend’s MegaMoE benchmark fuses expert-parallel dispatch, two grouped GEMMs, SwiGLU, and combine. Its reported configuration is EP8, top-k 6, one shared expert, and averages across eight ranks. For a 384-expert case with 16,384 tokens, the README reports 846.3 TFLOPS for one listed hidden/intermediate configuration and 103.3 GB/s of communication bandwidth for another listed case.

The NVIDIA comparison is DeepEP plus DeepGEMM’s MoE layouts and the surrounding NCCL/NVLink/RDMA environment. The architectural problem is similar, but the numbers cannot be transferred across vendors without matching token counts, expert routing, precision, rank count, and network conditions.

Model-specific kernels: MQA logits and mHC prenorm

The Ascend project also includes kernels that are easy to miss if the release is described only as “a GEMM port.” The DeepGEMM-Ascend README labels its MQA logits benchmark for the DeepSeek Lightning Indexer path. It lists FP8 and FP4 prefill and decode cases; FP4 decode is reported at 124.2 microseconds for the documented shape, versus 150.9 microseconds for FP8.

The same README labels its HC prenorm kernel for DeepSeek’s mHC module, or Manifold-Constrained Hyper-Connections. Listed memory bandwidth reaches 3,463 GB/s at M=8,192 for the documented N and K values. These figures show workload-specific tuning, not coverage of every model operator or serving path.

Framework and runtime: CANN and TorchNPU versus CUDA

Huawei’s TorchNPU repository describes TorchNPU as a PyTorch adapter for Ascend NPUs. Its feature list includes native and custom PyTorch APIs, FSDP2, DTensor, collective operations, graph capture, profiling, WatchDog monitoring, and memory-management features.

The conceptual NVIDIA map is PyTorch plus the CUDA runtime and libraries. The operational difference is substantial: an Ascend installation needs the matching CANN release, driver, firmware, Python version, PyTorch version, and TorchNPU version. The TorchNPU documentation’s example installs CANN 9.1.0, PyTorch 2.12.0, and torch-npu 2.12.0, while DeepGEMM-Ascend separately documents CANN 9.20. That version difference is a warning to follow each repository’s compatibility matrix rather than mixing commands from unrelated guides.

Huawei’s own CANN documentation describes CANN as the software layer connecting frameworks and Ascend hardware, including runtime and operator-development paths. In practice, CANN is closer to the platform foundation than to a single CUDA library. A PyTorch model may retain familiar Python syntax while still requiring Ascend-specific kernels, graph behavior, and debugging procedures.

What remains outside the release

DeepSeek’s original open-infrastructure index includes storage and system-level projects such as 3FS and Smallpond, and it describes DualPipe, EPLB, and an inference-system architecture. Those projects are important reference points, but the Ascend-specific kernel release should not be read as a complete port of all of them.

A production cluster still needs hardware provisioning, drivers and firmware, CANN installation, interconnect configuration, distributed runtime support, observability, checkpointing, failure recovery, and a serving or training orchestrator. Huawei Cloud’s DeepSeek deployment guide illustrates the infrastructure side with instances, networking, subnets, and security groups; those services are deployment prerequisites, not part of DeepGEMM-Ascend itself.

That boundary matters most for training. Public kernels can lower a bottleneck without proving that a frontier-scale training run is reproducible from the public repositories alone.

Who can use the stack today

User or organizationCan use it now?What is requiredPractical verdict
Team with Ascend 950 hardwareYes, for supported kernelsLinux environment, matching driver/firmware, CANN 9.20, TorchNPU, compiler, and compatible Python/PyTorch setupBest-fit early adopter
Huawei Cloud or enterprise operator with Ascend capacityPotentially yesA supported instance/cluster plus the exact software matrix and deployment expertiseViable for controlled evaluation and serving
Research lab with older Ascend hardwareNot automaticallyConfirm device support; the initial DeepGEMM-Ascend release was developed and validated on Ascend 950 seriesDo not assume 910B/910C compatibility
NVIDIA-only workstation ownerNo for the Ascend kernelsAscend hardware is a documented requirementUse the NVIDIA DeepSeek repositories instead
Ordinary PyTorch developer without accelerator accessNot in a meaningful runtime senseCan inspect code and study APIs, but cannot reproduce the hardware benchmarksDocumentation access is not execution access
Team seeking a turnkey frontier-training replacementNo public proof yetNeeds a complete cluster, systems integration, and production validation beyond the kernel repositoriesTreat as an infrastructure program, not a pip install

The practical boundary is also visible in the requirements: the software remains tied to Ascend 950 hardware, CANN, and torch_npu for the documented release. That is a narrower claim than general accelerator portability.

What the published numbers do—and do not—prove

The 99.8% dense-GEMM figure is useful evidence that the listed Ascend kernel can use the tested device efficiently for selected shapes. The MegaMoE tables add evidence that DeepSeek’s engineers addressed fused expert-parallel workloads rather than only an isolated matrix multiply.

Neither result answers the questions a procurement or training team ultimately has:

  • What is end-to-end tokens-per-second on the full model?
  • What is cost per token at the target batch size?
  • How stable are long runs and restarts?
  • Which operators fall back to less-optimized paths?
  • How do interconnect, memory, and power compare with the intended NVIDIA cluster?
  • Are the same results reproducible outside the original test environment?

DeepSeek has published serious Ascend kernel work with setup details and selected performance tables. The release lowers the software barrier for teams already inside the Ascend ecosystem, but the hardware and version barriers remain.

FAQ

Is DeepSeek Ascend infrastructure fully open source?

No. DeepGEMM-Ascend and related documentation are public, but the material does not provide one turnkey DeepSeek training cluster with all dependencies and operational recipes.

Can I run DeepGEMM-Ascend on an NVIDIA GPU?

No. It targets Ascend 950 hardware. NVIDIA users should use the NVIDIA-oriented DeepSeek projects.

Does this prove DeepSeek trains frontier models on Ascend?

No. It proves that DeepSeek published Ascend kernels and benchmarked selected workloads. It does not establish a public, end-to-end frontier-training reproduction.

Use the stack now if you already control supported Ascend capacity and can own the CANN/TorchNPU compatibility work. With NVIDIA hardware alone, treat the release as a technical reference, not a runnable backend.