Skip to content

feat(harness): return failed output checks to the Agent - #217

Merged
harlan-zw merged 2 commits into
mainfrom
feat/package-skill-run-findings
Oct 7, 2026
Merged

harlan-zw merged 2 commits into
mainfrom
feat/package-skill-run-findings

Conversation

@harlan-zw

Copy link
Copy Markdown
Collaborator

🔗 Linked issue

Related to harlan-zw/which-nuxt#13, unjs/unhead#1009, skilld-dev/retriv#21, nuxt-modules/robots#339, nuxt/scripts#950

❓ Type of change

  • 📖 Documentation
  • 🐞 Bug fix
  • 👌 Enhancement
  • ✨ New feature
  • 🧹 Chore
  • ⚠️ Breaking change

📚 Description

Five direct generate-package-skill runs, each followed by an independent review-skill run, shipped the pull requests above. Each review found errors or warnings in the generated Skill, and most findings came from prose claims the Skill never called examples. One run tested npm latest (1.x) while the branch published 2.0 betas. Every run hand-wrote a code-block extractor. The @unhead/vue Skill also shipped a description with quotes and % that read as broken. Nothing caught it, because the Harness checks output only after the Agent finishes, then ends the run.

The Harness now returns failed output checks to the Agent in the same session, up to two times, and requires the description to be one plain YAML line. generate-package-skill gains the run findings: an internal-package gate, the branch version rule, prose claims and failure inputs as examples, the Skillgen output shape, and clearer README and files rules. It also bundles scripts/extract-blocks.mjs. serve-fixture.mjs gains --header and records response headers. review-skill defines its three levels and tests the version the Skill names.

Architecture: runAgent sends failed output checks back to the Agent session for up to two repair turns before promotion; the generation and review Skills it loads gain new rules and an extract-blocks script

🤖 AI disclosure: Harlan Agent Kit modified this description. My AI open-source policy.

Five direct runs (@nuxtjs/robots, @nuxt/scripts, @unhead/vue, which-nuxt,
retriv) and their independent reviews found the same gaps. Every run
hand-wrote a code-block extractor. Most review findings came from prose
claims, which the Skill never called examples. One run tested npm latest
(1.x) while the branch published 2.0 betas. Node fetch dropped the Host
header the robots run needed. Reviewers had no definition of the three
finding levels.

Add extract-blocks.mjs, --header and response headers to serve-fixture,
an internal-package gate, the branch version rule, prose-claim and
failure-input checks, and defined review levels. The Harness requests now
state the Skillgen output shape and keep fixtures out of the output.
A run whose output failed a deterministic check ended as InvalidSkill,
and the session was destroyed before the Agent saw why. The Agent still
held its context and could have fixed the issue in one turn. The Harness
now sends the issues back in the same session, up to two times.

The @unhead/vue Skill shipped a frontmatter description with quotes and %
that read as broken. The check now requires one plain YAML line without
double quotes, backticks, or %, and the Skill says so.
@harlan-zw harlan-zw added the harlan-agent-review Approve automated work for the current issue state or pull request head commit. label Oct 7, 2026
@harlan-github-agent harlan-github-agent Bot added harlan-agent-review-required Pull request triage requires an adversarial Review for this head commit. and removed harlan-agent-review Approve automated work for the current issue state or pull request head commit. labels Oct 7, 2026
@harlan-zw
harlan-zw merged commit ac1c91d into main Oct 7, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

harlan-agent-review-required Pull request triage requires an adversarial Review for this head commit.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant