<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>模型量化 &#8211; mylogs.cn</title>
	<atom:link href="https://mylogs.cn/tag/%e6%a8%a1%e5%9e%8b%e9%87%8f%e5%8c%96/feed/" rel="self" type="application/rss+xml" />
	<link>https://mylogs.cn</link>
	<description>发现、记录、分享</description>
	<lastBuildDate>Mon, 03 Aug 2026 23:17:05 +0000</lastBuildDate>
	<language>zh-Hans</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	
	<item>
		<title>跑 Kimi 慢了一点，成本反降三成</title>
		<link>https://mylogs.cn/%e8%b7%91-kimi-%e6%85%a2%e4%ba%86%e4%b8%80%e7%82%b9%ef%bc%8c%e6%88%90%e6%9c%ac%e5%8f%8d%e9%99%8d%e4%b8%89%e6%88%90/</link>
					<comments>https://mylogs.cn/%e8%b7%91-kimi-%e6%85%a2%e4%ba%86%e4%b8%80%e7%82%b9%ef%bc%8c%e6%88%90%e6%9c%ac%e5%8f%8d%e9%99%8d%e4%b8%89%e6%88%90/#respond</comments>
		
		<dc:creator><![CDATA[steve, zhang]]></dc:creator>
		<pubDate>Mon, 03 Aug 2026 23:16:38 +0000</pubDate>
				<category><![CDATA[科技]]></category>
		<category><![CDATA[AI推理]]></category>
		<category><![CDATA[Cloudflare]]></category>
		<category><![CDATA[Kimi]]></category>
		<category><![CDATA[智谱GLM]]></category>
		<category><![CDATA[月之暗面]]></category>
		<category><![CDATA[模型量化]]></category>
		<guid isPermaLink="false">https://mylogs.cn/%e8%b7%91-kimi-%e6%85%a2%e4%ba%86%e4%b8%80%e7%82%b9%ef%bc%8c%e6%88%90%e6%9c%ac%e5%8f%8d%e9%99%8d%e4%b8%89%e6%88%90/</guid>

					<description><![CDATA[你在海外应用里用到的 AI，背后越来越可能是中国团队做的模型。Cloudflare 8 月 3 日发的一篇工程 [&#8230;]]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">你在海外应用里用到的 AI，背后越来越可能是中国团队做的模型。Cloudflare 8 月 3 日发的一篇工程博客里，被挑出来重点优化的两个模型，正是月之暗面的 Kimi K 系列和智谱的 GLM——文章开头就写明，这两家的模型「用起来非常舒服，但因为显存吃紧，非常难高效地跑起来」。</p>



<figure data-wp-context="{&quot;imageId&quot;:&quot;6a729b9b00c96&quot;}" data-wp-interactive="core/image" data-wp-key="6a729b9b00c96" class="wp-block-image size-large aligncenter wp-lightbox-container"><img fetchpriority="high" decoding="async" width="1200" height="628" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://mylogs.cn/wp-content/uploads/2026/08/e4172982401b6805a2ee880967282083_header.webp" alt="跑 Kimi 慢了一点，成本反降三成" class="wp-image-3448" style="max-width:100%;height:auto;" srcset="https://mylogs.cn/wp-content/uploads/2026/08/e4172982401b6805a2ee880967282083_header.webp 1200w, https://mylogs.cn/wp-content/uploads/2026/08/e4172982401b6805a2ee880967282083_header-300x157.webp 300w, https://mylogs.cn/wp-content/uploads/2026/08/e4172982401b6805a2ee880967282083_header-1024x536.webp 1024w, https://mylogs.cn/wp-content/uploads/2026/08/e4172982401b6805a2ee880967282083_header-768x402.webp 768w" sizes="(max-width: 1200px) 100vw, 1200px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button><figcaption class="wp-element-caption">图片来源：Cloudflare 官方博客</figcaption></figure>





<p class="wp-block-paragraph">三位工程师讲了三招。第一招，把模型用来记住上下文的「键值缓存」从 16 位精度压到 8 位，Kimi K2.6 能同时装下的上下文量从约 68.6 万词元翻到约 137 万。反常识的地方在这儿：单个请求反而更慢了（每秒 137 个词元降到 125），可同时能处理的请求数从 32 个上限顶到 64 个，峰值吞吐比原来高出 41%，每词元成本降约三成。第二招，把 GLM 5.2 的权重从 8 位浮点压成 4 位整数，模型文件从 705GB 缩到 421GB，单张显卡占用从约 88GB 降到 52GB，解码速度不降反升 55%。</p>



<p class="wp-block-paragraph">说白了，这是拿「单个用户慢一点」换「同时能服务的人多一倍」——云厂商算的是总账，不是单笔账。而两项常用测评的分数几乎没掉（94.24 对 94.09、89.11 对 89.04），说明这笔账划得来。第三招则是给共享缓存加了一层校验，防止上百个请求读串页，代价不到 1% 的吞吐。</p>



<p class="wp-block-paragraph">有意思的是，中国开源模型如今成了海外云厂商眼里「最难伺候、也最值得下功夫」的一类。如果要接一个大模型进自己的项目，你会先看效果还是先看单价？</p>



<p class="wp-block-paragraph">来源：Cloudflare 官方博客《Smaller, faster, safer: running Kimi and GLM at scale》，作者 Alex Reneau、Kevin Flansburg、Chi McIsaac，2026 年 8 月 3 日</p>

]]></content:encoded>
					
					<wfw:commentRss>https://mylogs.cn/%e8%b7%91-kimi-%e6%85%a2%e4%ba%86%e4%b8%80%e7%82%b9%ef%bc%8c%e6%88%90%e6%9c%ac%e5%8f%8d%e9%99%8d%e4%b8%89%e6%88%90/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
	</channel>
</rss>
