Repository navigation
feat(seo): index only Skills admitted from trending boards - #322
Merged
Merged
Conversation
Any made-up /gh/<owner>/<repo>/<skill> URL answered 200 with a thin page and a canonical to the homepage. Google filed 2,621 of them as soft 404s, and the open URL space wastes a small crawl budget. A missing Skill now answers 404 with noindex and no canonical. A page with no Skill never canonicalises to the homepage.
Experiment, gate 2026-11-11. Admission is additive and persisted in skill_trending_admissions, so a page never flips between index and noindex. The Skill API and the skills sitemap share one predicate, so the sitemap equals the indexable set.
Internal links to 301, 404 and 410 URLs waste crawl. Tag chips link to the final page, related lists skip Skills whose SKILL.md is gone, and a renamed repository hub keeps the robots value the API gave it. The 8 tag pages and 15 unaudited marketing pages leave the index during the experiment; the freeze audit records target query and bar per page.
Google finds pages through links first. Each trending board now lists the admitted Skills that left it as plain links with real page links, the boards and /skills join the skills sitemap when they render index, and the homepage links every period. X-Robots-Tag repeats the robots meta the page rendered, so the header and the HTML share one decision.
Collaborator
Author
|
Checked by hand, since CI cannot cover it. Local dev server, seeded D1, Googlebot user agent:
|
Experiment D links harlanzw.com posts to two skilld pages. Both need to be indexable whatever their quality score, so a probe row waives the score and nothing else. Renumber the migration to 0129, because the IndexNow PR (#320) takes 0128.
1 task done
An empty admitted set would noindex every Skill page, which Googlebot could fetch after a deploy or a failed task. While the table is empty the rule uses the pre-experiment decision and logs it once per isolate. The migration also seeds today's 65 admitted Skills and the gscdump probe, so the set starts populated.
nuxtseo-cli lives in a private repository, so skilld can never list it. nuxtjs-seo is public and in the production registry.
1 task done
Collaborator
Author
🤖 PENDING
|
`nuxi preview --port 5678` starts wrangler on 8787 for the Cloudflare preset, so Playwright timed out waiting for 5678. The Worker also read an empty D1 under .output, so every Skill API call failed with 500 and a missing Skill page rendered its error state with status 200. The script now applies local migrations and runs wrangler dev against the same local D1 that `pnpm dev` uses. A fresh checkout gets a migrated, empty D1.
The page sets no canonical for a missing Skill, but nuxt-seo-utils adds a self canonical to every page that is not a Nuxt error. That is harmless on a 404. The spec now fails only when a canonical names another URL, which is the homepage bug the PR fixes.
A transient API error rendered a real Skill page as 200 noindex, which could drop it from Google. The failed state now answers 503 with Retry-After and emits no robots directive, so Google retries.
This was referenced Sep 30, 2026
Contributor
🤖 MERGED
GitHub merged this pull request.
0b599464-d768-4eea-9a2c-342264c0f2a9 |
1 of 6 tasks
1 task done
harlan-zw
added a commit
that referenced
this pull request
Sep 30, 2026
Neither Skill sat on a trending board after #322, so both went noindex,follow. Replacements are admitted and index,follow in production.
harlan-zw
added a commit
that referenced
this pull request
Sep 30, 2026
Neither Skill sat on a trending board after #322, so both went noindex,follow and no longer matched their indexable partners.
1 task done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
❓ Type of change
📚 Description
Stacked on #317 (
fix/soft-404-missing-skills). The base is that branch, notmain.SEO experiment, gate 2026-11-11. Google holds skilld's curated URLs as "Discovered, currently not indexed" and crawls about 29 HTML pages a day. So only Skills that appeared on a trending board (week, month, all) stay indexable and in the skills sitemap. Every other Skill page is
noindex,followand leaves the sitemap.layers/registry/server/utils/trending-admission.tsnames the experiment, the gate, and the cull path.0129_skill_trending_admissions.sqlstores the admitted set withadmitted_at. An hourly task,admit-trending-skills, adds Skills and never removes one, so a page cannot flip between index and noindex.isSkillIndexableandSKILL_ADMITTED_SQL. The?supported=variant is gone, because it made the sitemap narrower than the pages.?page=links. The boards and/skillsjoin the skills sitemap when they renderindex. The homepage links Week, Month, and All time.X-Robots-Tagfrom the robots meta the page rendered. The robots module setindex, followbefore render, on pages whose HTML saidnoindex,follow.Size of the set (production D1, read only, 2026-09-30). Boards hold 95 unique Skills today (week 30, month 30, all 50). 65 of them pass the existing quality score and are not aggregators. That is the starting set, against 1,414 Skill URLs in the sitemap now. The set grows as boards move.
High-demand Skills against the rule (none added by hand):
seo_indexable = 0seo_indexable = 0in both reposProbe exceptions (experiment D, same gate, 2026-11-11). Two named pages stay indexable whatever their quality score, so harlanzw.com posts can link to them. They are rows with
first_board = 'probe', listed inPROBE_EXCEPTIONS. Cull path: empty that list andDELETE FROM skill_trending_admissions WHERE first_board = 'probe'./gh/harlan-zw/gscdumpis a single-Skill repository, so its hub URL is the Skill page. The registry holds it withseo_indexable = 0. The probe row admits it./gh/harlan-zw/nuxt-seo/nuxtjs-seoreplacesnuxtseo-cli, which lives in a private repository that skilld can never list.nuxtjs-seois public and in the production registry (the Skill API answers 200; its page is noindex today). The0129seed admits it as a probe row.Hygiene.
/gh/hyf0/vue-skills) shows "Source not found", so it isnoindexand leaves thesourcessitemap once the repo row stores the new GitHub identity. No 301: the registry keys Skills by the old name, so the new name has no page./skills/trending?page=Nhas its own canonical only for 1 < N <= the admitted list's page count. Page 1 canonicalises to the bare range. A page past the last answers 404, so old/skills/leaderboard?page=2..5links stop resolving./agents/*page the admission rule refuses, found from the pages folder./make-skillleaves the pages sitemap. The 8 tag pages go noindex and thetagssitemap is removed. That last one is a judgement call; revert[slug].get.tsand__sitemap__/tags.tsto undo it.Freeze audit (VISION principle 2).
layers/marketing/app/utils/page-admissions.tsrecords target query and bar. An audited page with no entry isnoindex,followand leaves the sitemap. Every/agents/*page is audited by default. The bar is 100 measured searches a month; I chose it, so adjust it if you disagree.Cull path for an admitted page: remove its entry.
📝 Migration
0129through the normal migration flow. It must merge after feat(indexnow): re-enable submission for the curated set #320 (IndexNow,0128_indexnow.sql). Renumber it if feat(indexnow): re-enable submission for the curated set #320 merges later. Until feat(indexnow): re-enable submission for the curated set #320 lands, the migration-naming test reports a gap at 0128, so CI fails on that one test.seo_indexabledecision and logstrending-admission / fallback-empty-setonce per isolate. In that state the sitemap follows the page rule, so it is wider than the old supported-only sitemap.DELETE FROM skill_trending_admissionsempties the set, and revertingisSkillIndexablerestores the old rule.