Skip to content

Skill verifier ignores tool sub-step evidence: most drafts rejected as resonance-low (often resonance=0), and the same policies re-crystallize every run #2460

Description

@chiefmojo

Pre-submission checklist

  • I have searched existing issues and this hasn't been mentioned before
  • I have read the project documentation and confirmed this issue doesn't already exist
  • This issue is specific to MemOS and not a general software issue

Bug Description

Summary

The skill verifier's evidence-resonance check (core/skill/verifier.ts, computeResonance) reads only userText, agentText and reflection from each evidence trace. The step extractor (core/capture/step-extractor.ts) writes tool sub-steps with agentText: "", and with userText: "" on every sub-step after the first. In tool-heavy sessions most evidence traces are such sub-steps. The resonance check sees empty strings for them, so it rejects reasonable drafts as resonance-low, often with resonance exactly 0. The content of those traces is in toolCalls and summary, and the resonance check never reads either.

Checked against main (a7367d07) and dev-v2.0.36 (568e9a1c): both have the same computeResonance and the same sub-step construction.

Root cause

// core/capture/step-extractor.ts — tool sub-steps
userText: i === 0 ? userText : "",
agentText: "",
toolCalls: [tc],

// core/skill/verifier.ts — computeResonance
for (const t of evidence) {
  const txt = `${t.userText}\n${t.agentText}\n${t.reflection ?? ""}`.toLowerCase();
  ...
  if (overlap >= 2) hit += 1;
}
return hit / evidence.length;   // must reach minResonance (0.5)

A sub-step trace contributes almost nothing to txt beyond its reflection, so it can't reach the 2-token overlap, and it still counts in the denominator. Any policy whose top-N evidence is mostly tool sub-steps is capped well below 0.5, however faithful the draft is.

Impact (measured on a 3-instance production deployment, Hermes adapter, 2026-10-05)

  • Success rate: skill runs succeed on about 2–4% of evaluated policies per day on every instance (skill.run.done: (crystallized + rebuilt) / evaluated, 15 days).
  • Rejection mix: about 90% of rejections are skill.verify.fail reason=resonance-low, and about 70% of those have resonance=0. Parse and validation failures are a small fraction.
  • Evidence makeup: on one instance, 88% of traces from the last 30 days have empty userText and agentText. All of them have toolCalls and a non-empty summary. Their reflection is usually a short classification label, which contributes no useful tokens.
  • Repeated LLM calls: nothing records a verify failure against the policy (skill.verification.failed is emitted with a placeholder skillId and no policyId). The same policies are therefore re-crystallized on every reward.updated run. One instance made 553 crystallize attempts on 71 distinct policies in 3 days. Each attempt is a full skill-evolver LLM call (~6k prompt and ~1.6k completion tokens), and the verifier then discards the result.

Reproduction

Read-only, nothing persisted. On three policies that had recently been rejected, I ran the deployed gatherEvidence → crystallizeDraft (live skill-evolver model) → verifyDraft:

policy (title) evidence traces traces with any user/agent text coverage resonance
"Execute Git Commands in Remote or Local Shell" 6 1 1.00 0.00
"Manage Skills" 6 2 1.00 0.33
"List files in a directory" 6 2 1.00 0.17

All three drafts described their evidence accurately (for example "run git commands with an explicit target directory, locally or over ssh"), and tool coverage was 1.00. What the drafts were compared against was mostly empty strings.

A minimal unit-level repro: a draft whose summary and steps mention apk add openssl-dev, with evidence made of two traces that each have userText: "", agentText: "" and toolCalls: [{ name: "shell", input: "apk add openssl-dev" }]. verifyDraft returns ok: false, resonance: 0.

Suggested fix

  1. Build each trace's resonance text from userText, agentText, summary, and the tool calls' names and inputs, in addition to the reflection text. Tool sub-steps then contribute what they actually did.
  2. Optionally, record verify failures against the policy, for example by emitting skill.verification.failed with the policyId. That makes a cooldown possible, so a policy that keeps failing with no new evidence doesn't cost an LLM call on every run.

Environment

  • memos-local-plugin v2.0.x (fork tracking main), Hermes adapter
  • skill evolver: deepseek-v4-flash family via an OpenAI-compatible provider (the model doesn't matter; the verifier is deterministic)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

area:pluginOpenClaw & Hermesstatus:needs-triageNeeds initial triage | 需要初步判断 & 问题复现

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions