<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>MI355X &#8211; mylogs.cn</title>
	<atom:link href="https://mylogs.cn/tag/mi355x/feed/" rel="self" type="application/rss+xml" />
	<link>https://mylogs.cn</link>
	<description>发现、记录、分享</description>
	<lastBuildDate>Tue, 29 Sep 2026 16:58:42 +0000</lastBuildDate>
	<language>zh-Hans</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	
	<item>
		<title>MLPerf 6.1：AMD 512 卡刷新推理纪录</title>
		<link>https://mylogs.cn/mlperf-6-1%ef%bc%9aamd-512-%e5%8d%a1%e5%88%b7%e6%96%b0%e6%8e%a8%e7%90%86%e7%ba%aa%e5%bd%95/</link>
		
		<dc:creator><![CDATA[steve, zhang]]></dc:creator>
		<pubDate>Thu, 17 Sep 2026 09:00:10 +0000</pubDate>
				<category><![CDATA[科技]]></category>
		<category><![CDATA[AI 推理]]></category>
		<category><![CDATA[AMD]]></category>
		<category><![CDATA[MI355X]]></category>
		<category><![CDATA[MLPerf]]></category>
		<category><![CDATA[Vera Rubin]]></category>
		<category><![CDATA[基准测试]]></category>
		<category><![CDATA[英伟达]]></category>
		<guid isPermaLink="false">https://mylogs.cn/mlperf-6-1%ef%bc%9aamd-512-%e5%8d%a1%e5%88%b7%e6%96%b0%e6%8e%a8%e7%90%86%e7%ba%aa%e5%bd%95/</guid>

					<description><![CDATA[MLCommons 近日发布 MLPerf Inference v6.1 基准测试套件。作为衡量不同硬件在 A [&#8230;]]]></description>
										<content:encoded><![CDATA[<figure data-wp-context="{&quot;imageId&quot;:&quot;6ac112b0a6390&quot;}" data-wp-interactive="core/image" data-wp-key="6ac112b0a6390" class="wp-block-image wp-lightbox-container"><img decoding="async" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" style="max-width:100%;height:auto" src="https://mylogs.cn/wp-content/uploads/2026/09/mlperf_v61_amd_mi355x_2026_header.webp" alt="MLPerf Inference v6.1 基准测试结果公布" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button><figcaption>图片来源：IT之家</figcaption></figure>
<p class="wp-block-paragraph">MLCommons 近日发布 MLPerf Inference v6.1 基准测试套件。作为衡量不同硬件在 AI 推理任务中吞吐与延迟表现的权威评测，最新 6.1 版本将重点放在推理吞吐、扩展效率，以及不同 GPU 配置下的实际表现。AMD、<a href="https://mylogs.cn/%e5%8f%b0%e7%a7%af%e7%94%b5-2nm-%e5%b9%b4%e5%ba%95%e6%9c%88%e4%ba%a7%e8%83%bd%e5%86%b2-12-%e4%b8%87%e7%89%87%ef%bc%8c%e6%8f%90%e5%89%8d%e4%b8%a4%e5%b9%b4%e8%be%be%e6%a0%87/">英伟达</a>、<a href="https://mylogs.cn/%e6%b7%b1%e8%93%9dl06%e8%bf%bd%e9%a3%8e%e7%89%8810%e6%9c%88%e4%b8%8a%e5%b8%82%ef%bc%8c%e9%85%8d%e7%a2%b3%e7%ba%a4%e7%bb%b4%e5%b0%be%e7%bf%bcai%e5%ba%a7%e8%88%b1/">英特尔</a>等厂商在结果公布后同步提交了各自的成绩单。</p>
<h2 class="wp-block-heading">AMD 派出的 512 卡答卷</h2>
<p class="wp-block-paragraph">AMD 以约 575 万个 token/秒的速度，在 512 个 GPU 的规模下创造了新的 MLPerf 提交纪录，这也是 AMD 迄今为止规模最大的参赛配置，意在证明其性能、规模、软件成熟度与可复现性。</p>
<p class="wp-block-paragraph">在 DeepSeek R1 这一<a href="https://mylogs.cn/chatgpt-pro-%e9%87%8d%e5%bc%80-200-%e7%be%8e%e5%85%83%e6%a1%a3%ef%bc%9aapi-%e4%bb%b7%e5%80%bc%e8%bf%91%e7%bf%bb%e5%80%8d/">大模型</a>负载测试中，AMD 派出 512 个 Instinct MI355X GPU 参赛。离线（强调批量处理能力，通常用于衡量系统在非实时场景下的最大吞吐）成绩为 2,901,950 tokens/s，服务器（强调在线服务场景中的吞吐与响应能力，更接近实际部署中的请求处理环境）成绩为 2,405,310 tokens/s。</p>
<h2 class="wp-block-heading">英伟达的 288 卡对照</h2>
<p class="wp-block-paragraph">同一项目中，英伟达提交了 GB300 与 GB200 的 288 GPU 配置。其中 GB300 离线成绩为 2,705,130 tokens/s，服务器成绩为 2,028,030 tokens/s。需要特别说明的是，两家提交的 GPU 数量并不相同——AMD 用 512 块，英伟达用 288 块，配置存在差异，因此以上数据仅供参考，不能直接当作同规模下的性能对比。</p>
<h2 class="wp-block-heading">Vera Rubin 首次现身</h2>
<p class="wp-block-paragraph">本次测试的一大看点是英伟达新一代架构 Vera Rubin（VR200）首次进入 MLPerf 推理基准，出现了 36 GPU 与 72 GPU 两种配置。在相同的 72 个 GPU 配置下，Vera Rubin（VR200）要比 GB300 快 95%；此外，其 36 GPU 配置版本甚至比 72-GPU 的 GB200 配置更快。</p>
<p class="wp-block-paragraph">在 GPT-OSS 120B 负载中，AMD 同样以史上最高的 512 GPU 配置在离线与服务器两类场景中领先；而在 72 GPU 与 8 GPU 段，英伟达的 GB300（Grace Blackwell Ultra）配置则占据上风。AMD 的 MI350P 在 8 GPU 配置下也快于上一代 MI300X。英特尔方面则带来了 Arc Pro B70 与至强（Xeon）的提交，英伟达的 RTX PRO 系列也在 8 GPU 与 4 GPU 配置下参与了测试。</p>
<h2 class="wp-block-heading">评测之外更值得看什么</h2>
<p class="wp-block-paragraph">从 512 卡到 288 卡，从已量产的 GB300 到首次亮相的 VR200，这一轮提交把 AI 推理吞吐的竞争推到了新的量级。对采购方而言，MLPerf 的价值不只在于“谁跑得更快”，更在于揭示不同规模、不同架构下的扩展效率——在真实业务里，往往不是单卡峰值，而是千卡协同能否不掉链子，才决定了一座 AI 工厂的真正产能。</p>]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
