<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Muse Voice Transcribe &#8211; mylogs.cn</title>
	<atom:link href="https://mylogs.cn/tag/muse-voice-transcribe/feed/" rel="self" type="application/rss+xml" />
	<link>https://mylogs.cn</link>
	<description>发现、记录、分享</description>
	<lastBuildDate>Tue, 01 Sep 2026 22:54:18 +0000</lastBuildDate>
	<language>zh-Hans</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	
	<item>
		<title>Meta 推出实时语音转写模型，70 多种语言即时成文</title>
		<link>https://mylogs.cn/meta-%e6%8e%a8%e5%87%ba%e5%ae%9e%e6%97%b6%e8%af%ad%e9%9f%b3%e8%bd%ac%e5%86%99%e6%a8%a1%e5%9e%8b%ef%bc%8c70-%e5%a4%9a%e7%a7%8d%e8%af%ad%e8%a8%80%e5%8d%b3%e6%97%b6%e6%88%90%e6%96%87/</link>
		
		<dc:creator><![CDATA[steve, zhang]]></dc:creator>
		<pubDate>Tue, 01 Sep 2026 22:54:18 +0000</pubDate>
				<category><![CDATA[科技]]></category>
		<category><![CDATA[AI 模型]]></category>
		<category><![CDATA[ASR]]></category>
		<category><![CDATA[Meta]]></category>
		<category><![CDATA[Muse Voice Transcribe]]></category>
		<category><![CDATA[实时语音识别]]></category>
		<category><![CDATA[语音转写]]></category>
		<guid isPermaLink="false">https://mylogs.cn/meta-%e6%8e%a8%e5%87%ba%e5%ae%9e%e6%97%b6%e8%af%ad%e9%9f%b3%e8%bd%ac%e5%86%99%e6%a8%a1%e5%9e%8b%ef%bc%8c70-%e5%a4%9a%e7%a7%8d%e8%af%ad%e8%a8%80%e5%8d%b3%e6%97%b6%e6%88%90%e6%96%87/</guid>

					<description><![CDATA[Meta 发布首款实时音频感知模型 Meta 于 9 月 1 日发布 Muse Voice Transcrib [&#8230;]]]></description>
										<content:encoded><![CDATA[<h2 class="wp-block-heading">Meta 发布首款实时音频感知模型</h2>

<p class="wp-block-paragraph">Meta 于 9 月 1 日发布 Muse Voice Transcribe，这是其第一款实时音频感知模型，由 Meta Superintelligence Labs 打造。该模型将流式语音识别、说话人分离与时间端点检测结合在一起，能够在说话的同时完成转写，并自动区分录音中 <strong>20 位以上</strong>的不同说话人，判断某人何时说完一句话，全程无需额外的后处理步骤。</p>

<figure data-wp-context="{&quot;imageId&quot;:&quot;6a978717ca990&quot;}" data-wp-interactive="core/image" data-wp-key="6a978717ca990" class="wp-block-image size-large wp-lightbox-container"><img fetchpriority="high" decoding="async" width="1200" height="628" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://mylogs.cn/wp-content/uploads/2026/09/inline_1788303252034.webp" alt="Meta AI Mac 应用界面" class="wp-image-5863" srcset="https://mylogs.cn/wp-content/uploads/2026/09/inline_1788303252034.webp 1200w, https://mylogs.cn/wp-content/uploads/2026/09/inline_1788303252034-300x157.webp 300w, https://mylogs.cn/wp-content/uploads/2026/09/inline_1788303252034-1024x536.webp 1024w, https://mylogs.cn/wp-content/uploads/2026/09/inline_1788303252034-768x402.webp 768w" sizes="(max-width: 1200px) 100vw, 1200px" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button><figcaption>图片来源：9to5Mac</figcaption></figure>

<h2 class="wp-block-heading">支持 70 多种语言，按小时计费</h2>

<p class="wp-block-paragraph">Muse Voice Transcribe 在超过 <strong>70 种语言</strong>上完成训练，发布时已有 25 种语言通过验证，可处理超过一小时的音频，并支持句内与句间的母语级代码切换。模型还引入了名为「自适应延迟」的机制：对容易识别的语音快速输出，对困难词汇则调用更多音频上下文后再做判断。</p>

<p class="wp-block-paragraph">据 Meta 介绍，该模型在 Artificial Analysis 的流式语音转写排行榜上排名首位（截至 9 月 1 日）。通过 Meta Model API 调用时，价格为每 1000 音频分钟 3 美元，约合每小时 0.18 美元。</p>

<h2 class="wp-block-heading">已接入 Mac 端与编程工具</h2>

<p class="wp-block-paragraph">目前，该模型已为 Meta AI for Mac 与 Muse Code 中的听写功能提供支持。在 Mac 上，用户按住 Fn 键即可向任意应用进行语音输入。相关能力由 Meta 研究员 Spencer Barnett 于 9 月 1 日在社交平台 X 上公布。</p>

<h2 class="wp-block-heading">中文市场用户需注意</h2>

<p class="wp-block-paragraph">需要说明的是，Muse Voice Transcribe 目前主要依托 Meta AI 生态，面向以英语为主的市场，在中国内地并未正式提供服务。对于中文用户，类似的实时语音转写与会议纪要能力，已可由国内多家语音技术厂商与办公协作产品提供；本文仅作技术动态介绍，不构成使用建议。</p>]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
