Skip to content

Support permanent redirects and dotted finding aid IDs - #125

Open
gkostin1966 wants to merge 10 commits into
mainfrom
ARC-192/dots-in-ids-in-urls
Open

gkostin1966 wants to merge 10 commits into
mainfrom
ARC-192/dots-in-ids-in-urls

Conversation

@gkostin1966

Copy link
Copy Markdown
Collaborator

Support permanent redirects and dotted finding aid IDs

Summary

This change adds permanent redirects for finding aids whose EAD IDs have changed and establishes one canonical identifier policy throughout the application.

Finding aid IDs may change for several reasons:

  • The legacy application replaced dots with hyphens when constructing Solr document IDs.
  • A stakeholder may assign a new EAD ID to an existing finding aid.
  • The same finding aid may be renamed more than once over its lifetime.

The application now treats the redirect map as a history of exact old-to-new ID changes and returns an HTTP 301 Moved Permanently response for obsolete IDs. Redirect chains are resolved directly to the newest ID so clients and search engines do not need to traverse multiple redirects.

Canonical identifier policy

Solr document IDs are now the canonical operational identifiers used by routes, downloads, generated artifacts, indexing jobs, and deletion jobs.

Canonical Solr IDs:

  • Preserve dots.
  • Strip surrounding whitespace.
  • Use lowercase characters.

For example:

Source EAD ID:  Umich-WCL-F-103.1dub
Solr ID:        umich-wcl-f-103.1dub

The original EAD ID remains available in ead_ssi as source metadata, including its original capitalization. Application behavior no longer depends on the raw ead_ssi value.

The custom UmArclight::NormalizedId class is loaded directly by the Traject configuration. This is important because some indexing commands execute Traject as a standalone process and do not load Rails initializers. Loading the normalizer from the Traject configuration ensures that all indexing paths produce the same IDs.

Component reference IDs are also lowercased during indexing so collection and component URLs follow the same canonical-ID policy.

Permanent redirects

config/redirect_map.rb stores lowercase old-to-new ID mappings. It supports both legacy normalization redirects and stakeholder-requested EAD ID changes.

For a rename history such as:

A -> B
B -> C
C -> D

requests for A, B, or C receive a single permanent redirect to D.

The redirect resolver:

  • Performs case-insensitive lookups.
  • Preserves dots.
  • Follows an arbitrary number of rename mappings.
  • Detects redirect cycles.
  • Returns the lowercase canonical target.
  • Redirects historical component URLs by resolving the mapped finding-aid root
    while preserving the component suffix.

An exact redirect-map entry takes precedence over component-prefix matching.
When multiple historical finding-aid IDs overlap, the resolver uses the
longest matching root.

Mixed-case requests for current IDs are also redirected to their lowercase canonical URL. This prevents case-sensitive Solr lookups from returning a 404 for an otherwise valid finding aid.

Redirect behavior applies to:

  • Finding aid and component show pages.
  • XML downloads.
  • HTML downloads.
  • PDF downloads.
  • Arclight hierarchy endpoints.

Existing query parameters are retained in redirected URLs.

Routing

Catalog routes now accept the complete non-slash path segment as :id. This prevents Rails from interpreting a dot in an ID as a format separator.

Known historical IDs and noncanonical mixed-case IDs are routed to catalog#permanent_id_redirect. Lowercase current IDs continue to use the normal Blacklight catalog#show route.

The redirect is implemented at the HTTP layer rather than through synthetic Solr documents. This provides a real 301 Moved Permanently response without introducing redirect-only records into search results, facets, sitemaps, or other Solr-backed features.

Artifact and deletion behavior

SolrDocument#finding_aid_id returns the root Solr ID:

  • For a collection, it returns the collection's own Solr ID.
  • For a component, it returns the component's _root_ collection ID.

XML, HTML, PDF, and temporary package filenames now use this canonical finding aid ID instead of the raw EAD ID.

Finding aid deletion now removes the complete nested Solr document block using _root_. The supplied ID is normalized and Solr-escaped before the delete query is sent.

Redirect map safeguards

The redirect map is frozen and has automated checks that:

  • Every key and value is lowercase.
  • No redirect cycles exist.

The resolver also detects cycles at runtime and raises an explicit error rather than looping indefinitely.

Sample data

Four Bentley sample EAD files with dotted IDs were added:

  • umich-bhl-87265.0
  • umich-bhl-87265.4
  • umich-bhl-87265.11
  • umich-bhl-87265.25

These provide representative data for verifying indexing, routing, and artifact behavior with dotted finding aid IDs.

Deployment considerations

This changes the canonical Solr ID format for records whose prior IDs contained dots or uppercase characters.

Deployment requires:

  1. Reindexing finding aids so Solr contains the new lowercase IDs with dots preserved.
  2. Regenerating or renaming existing XML, HTML, and PDF artifacts to use canonical Solr IDs.
  3. Removing stale documents indexed under the previous ID format if the reindex process does not clear the index first.

The redirect map preserves access to obsolete public URLs after the new index is deployed.

Test coverage

Coverage includes:

  • Lowercase, dot-preserving ID normalization.
  • Standalone Traject indexing behavior.
  • Preservation of the original EAD ID in ead_ssi.
  • Lowercase component IDs.
  • Redirect chains resolving directly to the final target.
  • Historical component redirects that preserve component suffixes.
  • Longest-prefix matching for overlapping historical finding-aid IDs.
  • Case-insensitive redirect lookup.
  • Mixed-case current-ID canonicalization.
  • Redirect cycle detection.
  • Lowercase redirect-map validation.
  • Dotted-ID route recognition.
  • Query parameter preservation.
  • XML download and hierarchy endpoint redirects.
  • Root finding aid IDs for collections and components.
  • Normalized and escaped nested-block deletion.

Replaced the Rails-only normalizer override with UmArclight::NormalizedId, loaded directly by standalone Traject and Rails ingestion.
Canonicalized collection and component Solr IDs to lowercase while preserving dots.
Added canonical redirects for mixed-case current IDs, historical IDs, downloads, and hierarchy routes.
Redirect chains resolve directly to the newest ID and preserve query parameters.
Normalized deletion inputs and retained safe nested-block deletion through root.
Removed the unrelated README change.
Expanded coverage for standalone normalization, mixed-case components, chained redirects, query parameters, downloads, hierarchy routes, map casing, and cycles.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant