Pre-submission checklist
Bug Description
Summary
The skill verifier's evidence-resonance check (core/skill/verifier.ts, computeResonance) reads only userText, agentText and reflection from each evidence trace. The step extractor (core/capture/step-extractor.ts) writes tool sub-steps with agentText: "", and with userText: "" on every sub-step after the first. In tool-heavy sessions most evidence traces are such sub-steps. The resonance check sees empty strings for them, so it rejects reasonable drafts as resonance-low, often with resonance exactly 0. The content of those traces is in toolCalls and summary, and the resonance check never reads either.
Checked against main (a7367d07) and dev-v2.0.36 (568e9a1c): both have the same computeResonance and the same sub-step construction.
Root cause
// core/capture/step-extractor.ts — tool sub-steps
userText: i === 0 ? userText : "",
agentText: "",
toolCalls: [tc],
// core/skill/verifier.ts — computeResonance
for (const t of evidence) {
const txt = `${t.userText}\n${t.agentText}\n${t.reflection ?? ""}`.toLowerCase();
...
if (overlap >= 2) hit += 1;
}
return hit / evidence.length; // must reach minResonance (0.5)
A sub-step trace contributes almost nothing to txt beyond its reflection, so it can't reach the 2-token overlap, and it still counts in the denominator. Any policy whose top-N evidence is mostly tool sub-steps is capped well below 0.5, however faithful the draft is.
Impact (measured on a 3-instance production deployment, Hermes adapter, 2026-10-05)
- Success rate: skill runs succeed on about 2–4% of evaluated policies per day on every instance (
skill.run.done: (crystallized + rebuilt) / evaluated, 15 days).
- Rejection mix: about 90% of rejections are
skill.verify.fail reason=resonance-low, and about 70% of those have resonance=0. Parse and validation failures are a small fraction.
- Evidence makeup: on one instance, 88% of traces from the last 30 days have empty
userText and agentText. All of them have toolCalls and a non-empty summary. Their reflection is usually a short classification label, which contributes no useful tokens.
- Repeated LLM calls: nothing records a verify failure against the policy (
skill.verification.failed is emitted with a placeholder skillId and no policyId). The same policies are therefore re-crystallized on every reward.updated run. One instance made 553 crystallize attempts on 71 distinct policies in 3 days. Each attempt is a full skill-evolver LLM call (~6k prompt and ~1.6k completion tokens), and the verifier then discards the result.
Reproduction
Read-only, nothing persisted. On three policies that had recently been rejected, I ran the deployed gatherEvidence → crystallizeDraft (live skill-evolver model) → verifyDraft:
| policy (title) |
evidence traces |
traces with any user/agent text |
coverage |
resonance |
| "Execute Git Commands in Remote or Local Shell" |
6 |
1 |
1.00 |
0.00 |
| "Manage Skills" |
6 |
2 |
1.00 |
0.33 |
| "List files in a directory" |
6 |
2 |
1.00 |
0.17 |
All three drafts described their evidence accurately (for example "run git commands with an explicit target directory, locally or over ssh"), and tool coverage was 1.00. What the drafts were compared against was mostly empty strings.
A minimal unit-level repro: a draft whose summary and steps mention apk add openssl-dev, with evidence made of two traces that each have userText: "", agentText: "" and toolCalls: [{ name: "shell", input: "apk add openssl-dev" }]. verifyDraft returns ok: false, resonance: 0.
Suggested fix
- Build each trace's resonance text from
userText, agentText, summary, and the tool calls' names and inputs, in addition to the reflection text. Tool sub-steps then contribute what they actually did.
- Optionally, record verify failures against the policy, for example by emitting
skill.verification.failed with the policyId. That makes a cooldown possible, so a policy that keeps failing with no new evidence doesn't cost an LLM call on every run.
Environment
- memos-local-plugin v2.0.x (fork tracking
main), Hermes adapter
- skill evolver:
deepseek-v4-flash family via an OpenAI-compatible provider (the model doesn't matter; the verifier is deterministic)
Pre-submission checklist
Bug Description
Summary
The skill verifier's evidence-resonance check (
core/skill/verifier.ts,computeResonance) reads onlyuserText,agentTextandreflectionfrom each evidence trace. The step extractor (core/capture/step-extractor.ts) writes tool sub-steps withagentText: "", and withuserText: ""on every sub-step after the first. In tool-heavy sessions most evidence traces are such sub-steps. The resonance check sees empty strings for them, so it rejects reasonable drafts asresonance-low, often with resonance exactly0. The content of those traces is intoolCallsandsummary, and the resonance check never reads either.Checked against
main(a7367d07) anddev-v2.0.36(568e9a1c): both have the samecomputeResonanceand the same sub-step construction.Root cause
A sub-step trace contributes almost nothing to
txtbeyond its reflection, so it can't reach the 2-token overlap, and it still counts in the denominator. Any policy whose top-N evidence is mostly tool sub-steps is capped well below 0.5, however faithful the draft is.Impact (measured on a 3-instance production deployment, Hermes adapter, 2026-10-05)
skill.run.done:(crystallized + rebuilt) / evaluated, 15 days).skill.verify.fail reason=resonance-low, and about 70% of those haveresonance=0. Parse and validation failures are a small fraction.userTextandagentText. All of them havetoolCallsand a non-emptysummary. Theirreflectionis usually a short classification label, which contributes no useful tokens.skill.verification.failedis emitted with a placeholderskillIdand nopolicyId). The same policies are therefore re-crystallized on everyreward.updatedrun. One instance made 553 crystallize attempts on 71 distinct policies in 3 days. Each attempt is a full skill-evolver LLM call (~6k prompt and ~1.6k completion tokens), and the verifier then discards the result.Reproduction
Read-only, nothing persisted. On three policies that had recently been rejected, I ran the deployed
gatherEvidence→crystallizeDraft(live skill-evolver model) →verifyDraft:All three drafts described their evidence accurately (for example "run git commands with an explicit target directory, locally or over ssh"), and tool coverage was 1.00. What the drafts were compared against was mostly empty strings.
A minimal unit-level repro: a draft whose summary and steps mention
apk add openssl-dev, with evidence made of two traces that each haveuserText: "",agentText: ""andtoolCalls: [{ name: "shell", input: "apk add openssl-dev" }].verifyDraftreturnsok: false,resonance: 0.Suggested fix
userText,agentText,summary, and the tool calls' names and inputs, in addition to the reflection text. Tool sub-steps then contribute what they actually did.skill.verification.failedwith thepolicyId. That makes a cooldown possible, so a policy that keeps failing with no new evidence doesn't cost an LLM call on every run.Environment
main), Hermes adapterdeepseek-v4-flashfamily via an OpenAI-compatible provider (the model doesn't matter; the verifier is deterministic)