Ascend NPU fork of nanochat for LLM training with torch_npu/HCCL (experimental)
-
Updated
Feb 24, 2026 - Python
Ascend NPU fork of nanochat for LLM training with torch_npu/HCCL (experimental)
Mini-SGLang port for Ascend NPUs with FIA, paged KV cache, ragged continuous batching, and request lifecycle safety.
Run nanochat training efficiently on Huawei Ascend NPUs with minimal code changes, supporting tokenizer, pretraining, and evaluation workflows.
Native AscendC Mamba2 selective scan / SSD forward-backward custom operator for Huawei Ascend 910B3 and 950PR, with CANN, torch_npu, A100 benchmarks and msprof profiling.
Add a description, image, and links to the torch-npu topic page so that developers can more easily learn about it.
To associate your repository with the torch-npu topic, visit your repo's landing page and select "manage topics."