Repository navigation
feat(seo): add a temporary sitemap of retired URLs - #325
Merged
Merged
Conversation
Any made-up /gh/<owner>/<repo>/<skill> URL answered 200 with a thin page and a canonical to the homepage. Google filed 2,621 of them as soft 404s, and the open URL space wastes a small crawl budget. A missing Skill now answers 404 with noindex and no canonical. A page with no Skill never canonicalises to the homepage.
Experiment E, remove 2026-11-11. Google still holds about 52k old URLs. Listing the ones that now answer 301, 404 or 410 with a fresh lastmod should get them recrawled and dropped sooner. Not submitted.
Contributor
🤖 MERGED
GitHub merged this pull request.
8fda238b-8322-4350-8ca8-b9cf2466713d |
1 task done
Content edits belonged to other work. The checker now fails on an empty or unloadable sitemap. lastmod is fixed because Google ignores one that changes on every run. After the removal date the handler skips D1.
`nuxi preview --port 5678` starts wrangler on 8787 for the Cloudflare preset, so Playwright timed out waiting for 5678. The Worker also read an empty D1 under .output, so every Skill API call failed with 500 and a missing Skill page rendered its error state with status 200. The script now applies local migrations and runs wrangler dev against the same local D1 that `pnpm dev` uses. A fresh checkout gets a migrated, empty D1.
The page sets no canonical for a missing Skill, but nuxt-seo-utils adds a self canonical to every page that is not a Nuxt error. That is harmless on a 404. The spec now fails only when a canonical names another URL, which is the homepage bug the PR fixes.
A transient API error rendered a real Skill page as 200 noindex, which could drop it from Google. The failed state now answers 503 with Retry-After and emits no robots directive, so Google retries.
2 tasks done
1 task done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
❓ Type of change
📚 Description
Stacked on #317 (
fix/soft-404-missing-skills). The base is that branch, notmain. It is independent of #322 and merges in any order.Experiment E from the death-zone plan, temporary, remove on 2026-11-11. Google still holds about 52k old URLs as "Crawled, currently not indexed", and it crawls about 29 HTML pages a day. A separate
retiredsitemap lists URLs that now answer 301, 404 or 410, each with one stablelastmod(2026-09-30, the experiment start). The aim is a faster recrawl, so Google drops them sooner. Mueller has hinted this works; treat it as practitioner advice, not documented policy. Scale rule: "Crawled, currently not indexed" falls by 10k in four weeks.I did not submit it to Search Console.
sitemap_index.xmllists it, so submit the index only if you want it read.Where the URLs come from (9,647 after de-duplication, checked against production on 2026-09-30):
/skillsanswer 200 and are left out)/skills/*, 345/orgs/*)source_resolved = 0, read from D1 on each requestMARKETING_REDIRECTS)The 8,806 are a dated snapshot in
data/retired-urls.json. The other three sources come from the same tables the middleware and redirects use, so they follow the site as it changes. I probed all 817 of those against production; each answered 301 or 410.scripts/check-retired-sitemap.tsre-probes the served sitemap. It exits non-zero if the sitemap fails to load, is empty, or any URL answers something other than 301, 404 or 410. Run it after deploy.Not covered:
/guides/*and/people/*(410) and the retired category redirects inrouteRules. None of them appear in Search Console pages, and no table lists them.After 2026-11-11 the handler returns an empty list without running its D1 queries. Removal steps are in the header of
layers/registry/server/utils/retired-sitemap.ts.The
lastmodstays fixed on purpose. Google ignores alastmodthat changes on every regeneration. This PR holds only the retired sitemap work; marketing and docs edits are gone.