HIVE-29813: (ACID) Improve the queries of the metrics collection and duplicate deletion by adding indexes - #6706
Open
kuczoram wants to merge 1 commit into
Open
HIVE-29813: (ACID) Improve the queries of the metrics collection and duplicate deletion by adding indexes#6706kuczoram wants to merge 1 commit into
kuczoram wants to merge 1 commit into
Conversation
…duplicate deletion by adding indexes
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



What changes were proposed in this pull request?
Adding some indexes to the HMS backend tables to improve the performance of the queries used by metrics collection and duplicate deletion.
The following indexes are proposed to improve the performance of the query in MetricsInfoHandler:
CREATE INDEX TXNS_STATE_TYPE_ID_IDX ON TXNS (TXN_STATE, TXN_TYPE, TXN_ID) ALGORITHM=INPLACE LOCK=NONE;CREATE INDEX HL_ACQUIRED_AT_IDX ON HIVE_LOCKS (HL_ACQUIRED_AT) ALGORITHM=INPLACE LOCK=NONE;CREATE INDEX CQ_STATE_COMMIT_TIME_IDX ON COMPACTION_QUEUE (CQ_STATE, CQ_COMMIT_TIME) ALGORITHM=INPLACE LOCK=NONE;For the duplicate deletion from the COMPLETED_TXN_COMPONENTS table, I propose to extend the existing index by adding the CTC_WRITEID and CTC_UPDATE_DELETE columns as well.
Why are the changes needed?
We faced some performance issues with these queries and trying to find a way to improve them. I did some testing with and without the proposed indexes and it can be seen that the indexes improve the performance of these queries. Please find the numbers in the Jira description.
Does this PR introduce any user-facing change?
No
How was this patch tested?
By unit tests and manually