<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
	<channel>
		<title>模型压缩 on 编程心语</title>
		<link>https://www.ithome.me/tags/%E6%A8%A1%E5%9E%8B%E5%8E%8B%E7%BC%A9/</link>
		<description>Recent content in 模型压缩 on 编程心语</description>
		<generator>Hugo</generator>
		<language>zh-CN</language>
		
		
		
		
			<lastBuildDate>Mon, 31 Aug 2026 11:00:00 +0800</lastBuildDate>
		
			<atom:link href="https://www.ithome.me/tags/%E6%A8%A1%E5%9E%8B%E5%8E%8B%E7%BC%A9/index.xml" rel="self" type="application/rss+xml" />
			<item>
				<title>端侧模型剪枝实战：剪掉不重要的参数，让模型更小更快</title>
				<link>https://www.ithome.me/post/2026/08/31/edge-model-pruning-guide/</link>
				<pubDate>Mon, 31 Aug 2026 08:00:00 +0800</pubDate>
				<guid>https://www.ithome.me/post/2026/08/31/edge-model-pruning-guide/</guid>
				<description>&lt;h2 id=&#34;压缩三件套还差最后一块拼图&#34;&gt;压缩三件套，还差最后一块拼图&lt;/h2&gt;&#xA;&lt;p&gt;前两周我们聊了端侧模型的&lt;strong&gt;量化&lt;/strong&gt;（INT8/INT4 把权重从 4 字节压到 0.5 字节）和&lt;strong&gt;知识蒸馏&lt;/strong&gt;（小模型学大模型的输出）。今天补上三件套的最后一块：&lt;strong&gt;剪枝（Pruning）&lt;/strong&gt;，把权重矩阵里不重要的参数直接置零。&lt;/p&gt;</description>
			</item>
			<item>
				<title>端侧模型知识蒸馏实战：把大模型的本事教给手机小模型</title>
				<link>https://www.ithome.me/post/2026/08/12/edge-ai-knowledge-distillation/</link>
				<pubDate>Wed, 12 Aug 2026 08:00:00 +0800</pubDate>
				<guid>https://www.ithome.me/post/2026/08/12/edge-ai-knowledge-distillation/</guid>
				<description>&lt;p&gt;上次我们用 INT8 量化把 MobileNetV2 从 14 MB 压到 3.5 MB（见&lt;a href=&#34;https://www.ithome.me/post/2026/08/07/on-device-quantization-model/&#34; rel=&#34;&#34;&gt;端侧模型 INT8 量化实战&lt;/a&gt;&#xA;）。但量化是事后减肥：模型该多大还多大，该多笨还多笨。今天聊另一种压缩思路——&lt;strong&gt;知识蒸馏（Knowledge Distillation）&lt;/strong&gt;，它发生在训练阶段：让小模型拜大模型为师，直接把大模型脑子里的判断力继承下来。&lt;/p&gt;&#xA;&lt;p&gt;一句话概括：&lt;strong&gt;量化管体积，蒸馏管智商&lt;/strong&gt;。两者不冲突，先蒸馏再量化，是端侧部署的标准组合拳。&lt;/p&gt;</description>
			</item>
	</channel>
</rss>
