Repository navigation
perf: parse records in linear time instead of quadratic - #649
Merged
Merged
Conversation
Parsing scaled with entries x document length, so large sitemaps became
unusable: a real 40,000-entry, 8.8MB product sitemap took about 229 seconds and
was killing a CPU-bounded Worker mid-request.
Two costs in findRecordClose ran over the whole remaining buffer on EVERY
record:
1. input.toLowerCase() copied the entire remainder per call, purely to do a
case-insensitive search. A sticky case-insensitive RegExp now scans in place.
2. indexOf('<!--') and indexOf('<![CDATA[') were unbounded. Ordinary sitemaps
contain neither, so both ran to the end of the document just to report
'absent', once per record. They are now bounded to [cursor, closeIndex), the
only window where a hidden section can change which close tag is real.
extractRecord had the same shape: source.toLowerCase().startsWith(...) copied
the remainder to compare a few characters. startsWithCI lowercases only the
prefix-length slice.
Behaviour is unchanged. The existing guards for close tags hidden inside CDATA
and for comments split across chunks still pass, along with the rest of the
suite (42 tests).
Measured on the same 8.8MB / 40,000-entry file: ~229s to ~0.55s. The new
regression test pins the shape rather than a wall-clock number: against the
previous code it fails at 6.2s for 8,000 entries and shows a 16.5x growth
factor for 4x the entries, which is the quadratic signature.
commit: |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
🔗 Linked issue
None. Found while debugging a production 503: a Cloudflare Worker was being killed mid-request parsing a large customer sitemap.
❓ Type of change
📚 Description
collectSitemap/parseSitemapscaled with entries × document length, so large sitemaps became unusable. A real 40,000-entry, 8.8MB product sitemap took about 229 seconds.Two costs in
findRecordCloseran over the whole remaining buffer on every record:input.toLowerCase()copied the entire remainder per call, purely to do a case-insensitive search. A sticky case-insensitiveRegExpnow scans in place.indexOf('<!--')andindexOf('<![CDATA[')were unbounded. Ordinary sitemaps contain neither, so both ran to the end of the document just to report "absent" — once per record. They are now bounded to[cursor, closeIndex), the only window where a hidden section can change which close tag is real.extractRecordhad the same shape:source.toLowerCase().startsWith(...)copied the remainder to compare a handful of characters.startsWithCIlowercases only the prefix-length slice.Measurements
Same file throughout (8.8MB, 40,000
<url>entries):Growth went from ~4x per doubling (quadratic) to ~2x (linear).
Behaviour is unchanged
This is purely a complexity fix. The existing guards that matter here still pass:
Full package suite green: 42 tests. ESLint clean.
The new test pins the shape, not a wall-clock number
test/parse-perf.test.tsasserts a large sitemap parses under a generous 5s ceiling, and that quadrupling the entry count grows runtime by less than 10x. Against the previous implementation it fails at 6.2s for 8,000 entries with a 16.5x growth factor for 4x the entries, which is the quadratic signature (4² = 16). The thresholds are deliberately loose so this separates linear from quadratic without flaking on a noisy CI machine.Note for downstream
nuxtseo.comis currently carrying this as apnpm patchonsitemapd@0.2.0to unblock production. Once this lands and is released, that patch should be dropped.