MUSHARE

Blog博客

Regression-testing a skill on a real AI assistant: from 155 seconds a case to 10 minutes a run在真实 AI 助手上做技能回归测试:从每个用例 155 秒到每轮 10 分钟

The problem

We build Mushare, a messaging network for AI assistants: you tell Meta's Muse "ask Alice whether Saturday or Sunday works", and Alice's Muse shows her a card she answers with a tap. On Muse, Mushare is a skill: a SKILL.md with references, plus a small Python CLI that Muse runs on its own cloud computer.

The CLI has 200 unit tests and they all pass. They say nothing about the part that fails most: whether Muse reads the skill the way we meant. Does it pick our skill for "tell Alice I'm late"? Does it run the right command? Does it show the card, or retell the card in its own words? Only a real assistant can answer that, and its answers vary from run to run.

So we built a regression suite that talks to Muse the way a user does, in the browser, and scores what comes back.

The harness

Each test account is a real Muse account, logged in once in its own Chrome profile. A Node script drives the browsers with Playwright over the Chrome DevTools Protocol. A test case is a sentence a real user would say, plus what should happen. For one case the script:

  1. Opens a new side chat, so no earlier conversation leaks in.
  2. Says the request, for example "Ask Alice: Costco on Saturday or Sunday? Give her two options", sometimes with a receipt image it drew for this run.
  3. Waits until Muse is done (more on that below).
  4. Collects the Mushare cards on the page (each card is an iframe our CLI renders), the reply text and a screenshot.
  5. After the run, pulls our server's tool-call log and matches each call to the case by account and time window.
  6. Scores each case: at least one card, the card contains the right words, the right server calls happened (send_message, recall_message), the content went out as a share card rather than plain text, and Muse did not say "I already sent this".

The output is a one-line-per-case summary (PASS, FAIL, FLAKY, with the reason and a ▲ or ▼ against the last run) and an HTML report with screenshots. We run only the affected cases while iterating on the skill text, and the full suite before each release.

Memory is state

The first surprise: the same SKILL.md scored 15/16 in the morning and 10/16 in the afternoon. Nothing had changed in our code. What had changed was Muse. After dozens of test runs a day, it remembered them. It answered "what do I need to deal with?" from memory instead of running our command. It told us "I already sent this receipt to Alice last night" and stopped. It had even stored a "preference" that the user wanted a hand-computed bill preview, learned from our own test sentence.

Three changes made runs comparable again:

  • Every case starts with one line asking Muse to pause its memory for that chat. The only cases that keep memory are the ones that test "look it up now, don't answer from memory".
  • Content is generated fresh for every run: random products to compare, a different restaurant and dishes on a receipt drawn as an image, times and reasons for each message. A repeated try within one run gets new content too.
  • A reply that says "already sent" or "send it again?" is scored as stale state, not as a skill failure.

The lesson is that an assistant's memory is test state, like a database. Reset it, or your results drift through the day.

The 155-second wait

A full run of 21 cases took 62 minutes, and almost every case took exactly 154 or 155 seconds, even "what can Mushare do?". Assistants are slow, but not that regular.

We watched the page with a second, read-only script and found the cause in our own code. To read the chat, the harness asked Playwright for the text of main. Muse's web app no longer had a main element. So every look waited for Playwright's 30-second timeout, then fell back to reading the whole page. "Done" meant three unchanged looks in a row, so every wait was about 155 seconds. Muse itself had usually finished in 30 to 40.

The fix had two parts:

  • Read the page text in place (document.body.innerText) and cut off the side panel, instead of waiting for an element.
  • Decide "done" from what the app shows while it works: a Stop button next to the input, and the assistant's status line ("Drafting message", "Sending message") instead of "Connected". Done is no Stop button, a Connected status and six quiet seconds.

The Stop button alone was not enough: between tool steps it disappears for a while, and one case stopped early while Muse was still working. The status line closed that gap.

A full run now takes 9 to 10 minutes. We also log, per wait, when the chat changed and what changed, so the next slow step shows up in the report instead of in our patience.

Two account groups, and what fresh accounts revealed

To halve the run time we added a second pair of test accounts and ran the two pairs side by side. One account per browser profile: tabs in one browser share a login. Cases name roles (the sender, the friend), not accounts, and a case that depends on another ("confirm the bill Leo sent") runs after it in the same group.

The new pair did more than save time. On the old accounts, "tell Alice I'll be late" always went through Mushare. On the new ones, Muse looked for Messenger and WhatsApp, said it had no way to reach Alice, and offered to connect Messenger. The old accounts had months of memory that Alice is a Mushare friend; a new user has none. Our scores had been measuring returning users only.

The cause was one line in the skill's frontmatter: includeInPrompt: false. With it, Muse does not list the skill among the skills it has on hand; it only finds it if it decides to search. Setting it to true puts the description line in every conversation (not the whole SKILL.md, just the description), and we moved the routing rule to the front of that line: if the user names a person and no app, check Mushare first. By the next full run the fresh accounts went from 5/10 to 8/10.

We now treat the two groups as two kinds of user: the old pair is a returning user, the new pair is a new one, and we compare each group with itself.

Ask the assistant why

When a case kept failing, the most useful debugging step was to open a new chat on that account and ask Muse why. It cannot see its earlier chats with memory paused, but it can describe how it decides, and that was usually enough:

  • "How are my Mushare reminders set?" got a description of Muse's own scheduled job instead of our card. Muse said it had never opened the skill: the question looked like something it could answer from its job list. We added that question to the routing rules in the description line.
  • "What do I need to deal with?" was answered from Muse's own task tracker. Muse pointed out a conflict we had written ourselves: the card rule said "add at most one line", so it would not put its own tasks after our card, and wrote everything as text instead. We allowed one exception.
  • "Send the school notice to Alice" went out as plain text, not a share card. Muse explained that the verb "send" mapped straight to our send command before it ever reached the sharing rules. The fix was to decide by the content first: content from a source, or with two or more facts, is a share.

One caution: the assistant's account of itself is a hypothesis, not a fact, so every fix still has to pass the cases.

Results

In two days the full suite went from 13/16 on one account pair to 22/22 on two, and a full run from about an hour to under ten minutes. Most of the score came from skill-text changes the suite made safe to try; one came from a test we had not written (a share rule that turned "confirm dinner at 7, add it to the calendar" into a share card instead of a calendar event, found while shooting screenshots and now its own case).

Skill versionWhat changedCases passedFull run
0.61.2Skill rewritten in Simplified Technical English13/16 (one pair)about 45 min
0.61.6Second account pair added12/2162 min
0.61.9Skill listed in every chat; routing rule first17/219 min (wait fix)
0.61.15General questions routed; send decides by content18/219 min
0.61.16Settling a time is a message with a calendar event22/22about 10 min

CI for the codebase went from 5.5 to 3 minutes the same week (three parallel shards), and releases that change only skill text no longer wait for it: CI cannot test how an assistant reads a sentence. The regression suite can.

What we would tell someone starting out

  • Test the skill where it runs. Unit tests cover your code; only the real assistant tells you how it reads your words.
  • Score from two sides. The page shows what the user saw; your server log shows what actually happened. A card without a send, or a send without a card, are both failures.
  • Treat memory as state. Pause it, generate fresh content, and label "already done" replies as stale instead of failing the skill.
  • Keep a fresh account. Your old test accounts are your most loyal users. A new one shows what a stranger gets.
  • Time every wait. A constant duration is a bug in your harness, not in the model.
  • Ask the agent, then verify. Its explanation is a good hypothesis; the cases decide.
  • Write the skill plainly. We write ours in ASD-STE100 Simplified Technical English: short sentences, one instruction each, the condition first. Muse told us it misses conditions buried mid-paragraph, and the scores agreed.

Rong Zhou is the CTO and one of two founders of Mushare (mushare.ai), a messaging network for AI assistants.

问题

我们在做 Mushare,一个 AI 助手之间的消息网络:你对 Meta 的 Muse 说"问问 Alice 周六还是周日方便",Alice 的 Muse 就给她看一张卡片,她点一下就回复了。在 Muse 上,Mushare 是一个技能:一份带参考文档的 SKILL.md,加上一个 Muse 在自己的云电脑上运行的小型 Python CLI。

这个 CLI 有 200 个单元测试,全部通过。但它们对最容易出问题的部分什么也说明不了:Muse 是否按我们的本意理解这个技能。听到"告诉 Alice 我要迟到",它会选我们的技能吗?会运行正确的命令吗?会把卡片展示出来,还是用自己的话把卡片复述一遍?只有真实的助手能回答这些问题,而且每次回答还不一样。

所以我们做了一套回归测试,像用户一样在浏览器里和 Muse 对话,再给返回的结果打分。

测试框架

每个测试账号都是一个真实的 Muse 账号,在各自的 Chrome 配置文件里登录一次。一个 Node 脚本通过 Chrome DevTools 协议用 Playwright 驱动这些浏览器。一个测试用例就是真实用户会说的一句话,加上应该发生的结果。对每个用例,脚本会:

  1. 打开一个新的侧边对话,避免之前的对话混进来。
  2. 说出请求,例如"Ask Alice: Costco on Saturday or Sunday? Give her two options",有时附上为这次运行画的一张收据图片。
  3. 等 Muse 做完(详见下文)。
  4. 收集页面上的 Mushare 卡片(每张卡片是我们 CLI 渲染的一个 iframe)、回复文字和一张截图。
  5. 运行结束后,拉取我们服务器的工具调用日志,按账号和时间窗口把每次调用对应到用例。
  6. 给每个用例打分:至少有一张卡片,卡片里有正确的词,服务器上发生了正确的调用(send_message、recall_message),内容是作为分享卡片而不是纯文字发出的,而且 Muse 没有说"我已经发过了"。

输出是每个用例一行的汇总(PASS、FAIL、FLAKY,附原因,以及和上次相比的 ▲ 或 ▼),外加一份带截图的 HTML 报告。修改技能文字时只跑受影响的用例,每次发布前跑全套。

记忆就是状态

第一个意外:同一份 SKILL.md,上午得 15/16,下午得 10/16。我们的代码一点没变,变的是 Muse。一天跑几十次测试之后,它记住了这些测试。它直接凭记忆回答"我有什么需要处理的?",而不运行我们的命令。它告诉我们"昨晚已经把这张收据发给 Alice 了",然后就停了。它甚至从我们自己的测试句子里学到并存下了一条"偏好":用户想要一份手算的账单预览。

三项改动让各次运行重新可以比较:

  • 每个用例都先用一句话请 Muse 在这个对话里暂停记忆。只有专门测试"现在去查,不要凭记忆回答"的用例才保留记忆。
  • 每次运行都重新生成内容:随机的比价商品,收据图片上换一家餐厅和菜品,每条消息的时间和理由都不同。同一次运行里的重试也换新内容。
  • 回复里出现"已经发过了"或"要再发一次吗?",记为状态残留,而不算技能失败。

教训是:助手的记忆和数据库一样,是测试状态。要么重置,要么结果会随着一天的推移而漂移。

155 秒的等待

21 个用例的全套运行要 62 分钟,而且几乎每个用例都恰好用了 154 或 155 秒,连"Mushare 能做什么?"也一样。助手是慢,但不会这么整齐。

我们用另一个只读脚本观察页面,在自己的代码里找到了原因。为了读取对话,测试框架向 Playwright 要 main 元素的文字,而 Muse 的网页版已经没有 main 元素了。于是每看一次都要等满 Playwright 的 30 秒超时,然后才退回去读整个页面。"完成"的标准是连续三次看到的内容不变,所以每次等待都是 155 秒左右,而 Muse 本身通常 30 到 40 秒就完成了。

修复分两部分:

  • 直接读取页面文字(document.body.innerText)并去掉侧边面板,而不是等某个元素出现。
  • 根据应用工作时显示的东西判断"完成":输入框旁边的停止按钮,以及助手的状态行("Drafting message"、"Sending message",而不是"Connected")。完成就是没有停止按钮、状态为 Connected、并且安静六秒。

只看停止按钮还不够:在两个工具步骤之间,它会消失一会儿,有一个用例就在 Muse 还在工作时提前结束了。状态行补上了这个缺口。

现在全套运行只要 9 到 10 分钟。我们还按每次等待记录对话何时变化、变了什么,这样下一个慢步骤会出现在报告里,而不是靠我们的耐心去发现。

两组账号,以及新账号暴露的问题

为了把运行时间减半,我们又加了一对测试账号,两对并行跑。每个浏览器配置文件一个账号:同一个浏览器里的标签页共用一个登录。用例写的是角色(发送方、朋友),不是具体账号;依赖另一个用例的用例("确认 Leo 发来的账单")在同一组里排在它后面运行。

新的一对不只是省了时间。在老账号上,"告诉 Alice 我要迟到"总是走 Mushare;在新账号上,Muse 去找 Messenger 和 WhatsApp,说它联系不上 Alice,还提议连接 Messenger。老账号有几个月的记忆,知道 Alice 是 Mushare 好友;新用户没有。我们的分数一直只衡量了老用户。

原因是技能 frontmatter 里的一行:includeInPrompt: false。这样设置时,Muse 不会把这个技能列在它手头的技能里,只有它决定去搜索时才找得到。改成 true 之后,描述行会出现在每一次对话里(不是整份 SKILL.md,只是描述),我们又把路由规则移到这一行的最前面:如果用户提到一个人而没有提应用,先查 Mushare。到下一次全套运行,新账号从 5/10 提高到了 8/10。

现在我们把两组当作两类用户:老的一对是回头用户,新的一对是新用户,每组只和自己比较。

问助手为什么

一个用例反复失败时,最有用的调试步骤是在那个账号上开一个新对话,直接问 Muse 为什么。暂停记忆时它看不到之前的对话,但它能描述自己是怎么做决定的,这通常就够了:

  • "我的 Mushare 提醒是怎么设置的?"得到的是 Muse 自己的定时任务的说明,而不是我们的卡片。Muse 说它根本没有打开这个技能:这个问题看起来它用自己的任务列表就能回答。我们把这个问题加进了描述行里的路由规则。
  • "我有什么需要处理的?"是用 Muse 自己的任务追踪来回答的。Muse 指出了一个我们自己写出来的矛盾:卡片规则说"最多加一行",所以它不会在我们的卡片后面列出自己的任务,干脆全部写成了文字。我们加了一条例外。
  • "把学校通知发给 Alice"作为纯文字发出,而不是分享卡片。Muse 解释说,动词"发"在还没读到分享规则之前就直接对应到了我们的 send 命令。修复办法是先看内容:来自某个来源的内容,或者包含两条以上事实的内容,就是分享。

需要提醒一点:助手对自己的解释是假设,不是事实,所以每个修复仍然要通过用例。

结果

两天里,全套测试从一对账号上的 13/16 提高到两对账号上的 22/22,全套运行从大约一小时缩短到不到十分钟。大部分分数来自技能文字的修改,回归测试让这些修改可以放心尝试;有一项来自一个我们没写过的测试(一条分享规则把"确认晚上 7 点吃饭,加到日历里"变成了分享卡片,而不是日历事件;这是在拍截图时发现的,现在它有了自己的用例)。

技能版本改了什么通过用例全套运行
0.61.2用简化技术英语重写技能13/16(一对账号)约 45 分钟
0.61.6加入第二对账号12/2162 分钟
0.61.9技能在每次对话里列出;路由规则放在最前17/219 分钟(修复等待)
0.61.15一般性问题也路由;发送按内容决定18/219 分钟
0.61.16约定时间是一条带日历事件的消息22/22约 10 分钟

同一周,代码库的 CI 从 5.5 分钟缩短到 3 分钟(三个并行分片),而且只改技能文字的发布不再等 CI:CI 测不了助手怎么理解一句话,回归测试可以。

给刚起步的人的建议

  • 在技能真正运行的地方测试。 单元测试覆盖你的代码;只有真实的助手能告诉你它怎么理解你的文字。
  • 从两边打分。 页面显示用户看到了什么;服务器日志显示实际发生了什么。有卡片没发送,或者发送了没卡片,都是失败。
  • 把记忆当作状态。 暂停记忆,生成新内容,把"已经做过了"的回复标为状态残留,而不是判技能失败。
  • 留一个新账号。 老测试账号是你最忠实的用户。新账号才能告诉你陌生人会得到什么。
  • 给每次等待计时。 恒定的耗时是测试框架的 bug,不是模型的问题。
  • 先问智能体,再验证。 它的解释是很好的假设;结论由用例来定。
  • 把技能写得简单直白。 我们用 ASD-STE100 简化技术英语写技能:句子短,每句一条指令,条件放在前面。Muse 告诉我们,它会漏掉埋在段落中间的条件,分数也印证了这一点。

Rong Zhou 是 Mushare(mushare.ai,AI 助手之间的消息网络)的 CTO,也是两位创始人之一。

Today setup asks you to paste a key once. One-tap install in the Muse app is on its way: get one email when it is live.

现在安装要粘贴一次密钥;Muse App 里的一键安装即将上线。上线时我们给你发一封邮件。