目录

ExecuTorch 端侧部署实战:从 torch.export 到 .pte 的完整工作流

做端侧部署,为什么多了一个 PyTorch 原生选项

做端侧 AI,最痛的一环往往是模型转换:PyTorch 训练完,先导出 ONNX,再转 TFLite 或 Core ML,算子一缺就回头改模型;量化、Profile、调试也是每套工具配一遍。

ExecuTorch 是 PyTorch 官方的端侧推理栈,思路不同:直接从 torch.export 出发,在同一套 PyTorch 工作流里完成后端特化(partitioner)与端侧运行,不需要中间格式,导出产物还能保留 PyTorch 的程序元数据和源码映射。它已经在 Instagram、WhatsApp、Facebook、Messenger 以及 Meta Quest、Ray-Ban Meta 设备上承载了数亿用户的功能,2026 年还发了 MLSys 论文(arXiv:2605.08195)。

本文跑通它的核心链路:把 PyTorch 模型导出成 .pte 文件并部署到端侧。内容截至 2026-09-17 验证,对应最新稳定版 v1.5.0(2026-09-16 刚发布)。

工作流四步

ExecuTorch 的部署链路可以概括为四步:

  1. 捕获:用 torch.export.export() 把 nn.Module 导出为可序列化的计算图
  2. 降级:to_edge_transform_and_lower() 传入后端 partitioner,把支持的子图委托给 CPU / NPU / DSP 后端,其余保留可移植 CPU 算子兜底
  3. 序列化:.to_executorch() 生成后端专属的 .pte 文件
  4. 运行:C++ / Python / Swift / Kotlin / WebAssembly 运行时加载 .pte 执行

注意第 3 步:.pte 是后端专属的。为 XNNPACK 导出的 .pte 不能直接拿到 Core ML 或 Qualcomm 后端跑,每个目标硬件需要单独导出。

环境准备

1
2
# 需 Python 3.10 至 3.14
pip install executorch

预编译 wheel 覆盖 Linux x86-64 / AArch64、macOS arm64、Windows x86-64;Android 可从 Maven Central 拉 AAR,Apple 平台通过 Swift Package Manager 集成。

实战一:导出视觉模型到 XNNPACK

从官方快速入门的经典例子看最小链路(导出一个加法模型验证环境):

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
import torch
from executorch.exir import to_edge_transform_and_lower
from executorch.backends.xnnpack.partition.xnnpack_partitioner import XnnpackPartitioner
from executorch.runtime import Runtime

class Add(torch.nn.Module):
    def forward(self, x, y):
        return x + y

model = Add().eval()
sample_inputs = (torch.ones(1), torch.ones(1))

et_program = to_edge_transform_and_lower(
    torch.export.export(model, sample_inputs),
    partitioner=[XnnpackPartitioner()]
).to_executorch()

with open("add.pte", "wb") as f:
    f.write(et_program.buffer)

runtime = Runtime.get()
method = runtime.load_program("add.pte").load_method("forward")
output = method.execute(sample_inputs)[0]
print("Output:", output)  # tensor([2.])

换成真实视觉模型,只改模型和输入。以 MobileNetV2 为例:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
import torch
from torchvision.models import mobilenet_v2, MobileNet_V2_Weights
from executorch.exir import to_edge_transform_and_lower
from executorch.backends.xnnpack.partition.xnnpack_partitioner import XnnpackPartitioner

model = mobilenet_v2(weights=MobileNet_V2_Weights.IMAGENET1K_V2).eval()
sample_inputs = (torch.randn(1, 3, 224, 224),)

et_program = to_edge_transform_and_lower(
    torch.export.export(model, sample_inputs),
    partitioner=[XnnpackPartitioner()],
).to_executorch()

with open("mobilenet_v2.pte", "wb") as f:
    f.write(et_program.buffer)

跑完得到 mobilenet_v2.pte,这就是部署产物。官方没有给出各机型的统一耗时基准,内存与速度需在目标设备上用 ETDump 自行实测。

实战二:用 export_llm 导出大模型

端侧 LLM 场景,手写 torch.export 加量化太繁琐,官方提供了高层 API export_llm,一条命令完成导出、量化与后端特化。当前官方支持的模型族包括 Llama 2/3/3.1/3.2、Qwen 2.5/3、Phi 3.5/4-mini、SmolLM2。

以 Llama 3.2 1B 为例,先准备配置文件:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
# config.yaml
base:
  model_class: llama3_2
  checkpoint: path/to/consolidated.00.pth
  params: path/to/params.json
  metadata: '{"get_bos_id":128000, "get_eos_ids":[128009, 128001]}'
model:
  use_kv_cache: True
  use_sdpa_with_kv_cache: True
quantization:
  embedding_quantize: 4,32
  qmode: 8da4w
backend:
  xnnpack:
    enabled: True
    extended_ops: True

其中 use_kv_cache 与 use_sdpa_with_kv_cache 是官方推荐任意 LLM 都打开的优化;8da4w 表示 int8 动态激活加 int4 权重量化,是 XNNPACK 后端的推荐量化档位。Llama 系列需要手动下载 consolidated.00.pth 与 params.json(Hugging Face 上 meta-llama/Llama-3.2-1B 的 original 目录有),其余模型族会自动从 Hugging Face 下载。

执行导出:

1
python -m executorch.extension.llm.export.export_llm --config config.yaml

想看哪些算子被委托给了 XNNPACK、哪些还留在 CPU,加两行配置:

1
2
3
debug:
  verbose: True
  generate_etrecord: True

verbose 会在日志里打出完整的委托情况表(委托子图数、委托节点数、逐算子分布),generate_etrecord 生成 ETRecord,可以用 ExecuTorch 开发者工具把算子耗时_trace_回源码行。顺带一提,跑 LLM 导出如果报 ModuleNotFoundError: No module named 'pytorch_tokenizers',补装 pip install pytorch-tokenizers 即可。

量化路径要和后端匹配:XNNPACK 走 TorchAO 量化(qmode),QNN / Core ML / Vulkan 走 pt2e 图级量化,两者不能混用。

后端怎么选

平台 可用后端
iOS / iPadOS XNNPACK(CPU)、Core ML、MLX(设备上,实验性)
Android XNNPACK、Vulkan(GPU)、Qualcomm、MediaTek、Arm VGF、Samsung Exynos
macOS / 桌面 XNNPACK、Core ML、MLX(实验性)、Metal/AOTInductor、WebGPU(实验性)
嵌入式 / MCU Arm Cortex-M + CMSIS-NN(beta)、Ethos-U、NXP eIQ Neutron、Cadence DSP、Zephyr、Arduino

iOS 上除了 XNNPACK,可以指定 Core ML 或 MLX partitioner 单独导出,把算子卸载到 Apple Neural Engine 或统一内存架构上。

iOS 集成

Apple 平台通过 Swift Package Manager 引入,仓库根目录自带 Package.swift,在 Xcode 里加 https://github.com/pytorch/executorch 即可。v1.5.0 起把 Core ML 与 MLX 的可链接库直接打包进了 Apple 框架与 SwiftPM 产物,还支持 ETDump 性能采集,端侧调性能不用再切回 Python。

v1.5.0 值得关注的新东西

v1.5.0(2026-09-16)这版更新很实:

  • C++ SDK 打进 wheel:CUDA、Core ML、MLX、OpenVINO、Qualcomm 各 delegate 都有了可链接库,还打包了 TorchAO 量化算子库,C++ 侧集成不用再从源码编译
  • LLM 服务化能力:多方法导出(multi-method)、批量请求调度、可取消执行、离图 KV-cache 布局,新增 Qwen3.5 MoE、Muse Glimmer、Supertonic、Voxtral 等工作流
  • Core ML / MLX:MLX 离图 KV-cache 的 flat / ring / cell 布局与共享池,融合注意力算子支持;Core ML 改进了对 tied embedding 和因果注意力掩码的处理
  • 可靠性:.pte / .ptd 加了 schema 版本校验,运行时兼容性政策保证稳定 API 导出的 .pte 至少能被下一个非 patch 版本加载执行

踩坑清单

  • .pte 不跨后端:换硬件 delegate 必须重新导出,不是改个加载参数就行
  • 算子覆盖率因后端而异:官方模型列表是起点不是兼容性清单,新架构可以走自定义 LLM 导出指南
  • 量化路径与后端绑定:TorchAO 给 XNNPACK,pt2e 给 QNN / Core ML / Vulkan
  • 端侧性能别只看 Python 侧数字:Python 运行时只用来验证正确性,性能必须在目标设备上重新测
  • 版本搭配:nightly wheel 不声明 torch 依赖,需要手动指定匹配的 PyTorch nightly

开源方案速查

  • ExecuTorch(pytorch/executorch):PyTorch 原生端侧部署,适合已在 PyTorch 生态、要多硬件后端的项目
  • Optimum ExecuTorch:Hugging Face 的测试好的导出配方,想省配置直接用它
  • TorchAO:PyTorch 量化算法库,ExecuTorch 的量化能力底层就是它

如果你之前一直在 PyTorch 和 ONNX / TFLite 之间来回转换,ExecuTorch 值得试一次:同一份模型代码、同一套导出工具链,覆盖从手机到 MCU 的部署目标。