Skip to content

perf: parse records in linear time instead of quadratic - #649

Merged
harlan-zw merged 1 commit into
mainfrom
fix/sitemapd-quadratic-record-scan
Aug 3, 2026
Merged

harlan-zw merged 1 commit into
mainfrom
fix/sitemapd-quadratic-record-scan

Conversation

@harlan-zw

Copy link
Copy Markdown
Collaborator

🔗 Linked issue

None. Found while debugging a production 503: a Cloudflare Worker was being killed mid-request parsing a large customer sitemap.

❓ Type of change

  • 📖 Documentation
  • 🐞 Bug fix
  • 👌 Enhancement
  • ✨ New feature
  • 🧹 Chore
  • ⚠️ Breaking change

📚 Description

collectSitemap / parseSitemap scaled with entries × document length, so large sitemaps became unusable. A real 40,000-entry, 8.8MB product sitemap took about 229 seconds.

Two costs in findRecordClose ran over the whole remaining buffer on every record:

  1. input.toLowerCase() copied the entire remainder per call, purely to do a case-insensitive search. A sticky case-insensitive RegExp now scans in place.
  2. indexOf('<!--') and indexOf('<![CDATA[') were unbounded. Ordinary sitemaps contain neither, so both ran to the end of the document just to report "absent" — once per record. They are now bounded to [cursor, closeIndex), the only window where a hidden section can change which close tag is real.

extractRecord had the same shape: source.toLowerCase().startsWith(...) copied the remainder to compare a handful of characters. startsWithCI lowercases only the prefix-length slice.

Measurements

Same file throughout (8.8MB, 40,000 <url> entries):

before after
8.8MB / 40k entries ~229,000ms ~550ms
0.19MB / 4k entries 261ms 40ms
0.37MB / 8k entries 954ms 77ms
0.75MB / 16k entries 3,724ms 139ms

Growth went from ~4x per doubling (quadratic) to ~2x (linear).

Behaviour is unchanged

This is purely a complexity fix. The existing guards that matter here still pass:

  • "does not treat closing markup inside CDATA as the record close"
  • "waits for comments split across chunks"

Full package suite green: 42 tests. ESLint clean.

The new test pins the shape, not a wall-clock number

test/parse-perf.test.ts asserts a large sitemap parses under a generous 5s ceiling, and that quadrupling the entry count grows runtime by less than 10x. Against the previous implementation it fails at 6.2s for 8,000 entries with a 16.5x growth factor for 4x the entries, which is the quadratic signature (4² = 16). The thresholds are deliberately loose so this separates linear from quadratic without flaking on a noisy CI machine.

Note for downstream

nuxtseo.com is currently carrying this as a pnpm patch on sitemapd@0.2.0 to unblock production. Once this lands and is released, that patch should be dropped.

Parsing scaled with entries x document length, so large sitemaps became
unusable: a real 40,000-entry, 8.8MB product sitemap took about 229 seconds and
was killing a CPU-bounded Worker mid-request.

Two costs in findRecordClose ran over the whole remaining buffer on EVERY
record:

1. input.toLowerCase() copied the entire remainder per call, purely to do a
   case-insensitive search. A sticky case-insensitive RegExp now scans in place.
2. indexOf('<!--') and indexOf('<![CDATA[') were unbounded. Ordinary sitemaps
   contain neither, so both ran to the end of the document just to report
   'absent', once per record. They are now bounded to [cursor, closeIndex), the
   only window where a hidden section can change which close tag is real.

extractRecord had the same shape: source.toLowerCase().startsWith(...) copied
the remainder to compare a few characters. startsWithCI lowercases only the
prefix-length slice.

Behaviour is unchanged. The existing guards for close tags hidden inside CDATA
and for comments split across chunks still pass, along with the rest of the
suite (42 tests).

Measured on the same 8.8MB / 40,000-entry file: ~229s to ~0.55s. The new
regression test pins the shape rather than a wall-clock number: against the
previous code it fails at 6.2s for 8,000 entries and shows a 16.5x growth
factor for 4x the entries, which is the quadratic signature.
@pkg-pr-new

pkg-pr-new Bot commented Aug 3, 2026

Copy link
Copy Markdown

Open in StackBlitz

npm i https://pkg.pr.new/@nuxtjs/sitemap@649

commit: e28be30

@harlan-zw harlan-zw changed the title fix(sitemapd): parse records in linear time instead of quadratic perf: parse records in linear time instead of quadratic Aug 3, 2026
@harlan-zw
harlan-zw merged commit 42fa8c2 into main Aug 3, 2026
10 checks passed
@harlan-zw
harlan-zw deleted the fix/sitemapd-quadratic-record-scan branch August 3, 2026 10:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant