Skip to content

Python: fix: prevent superlinear history growth by deduplicating messages in save_messages#7242

Open
PratikWayase wants to merge 1 commit into
microsoft:mainfrom
PratikWayase:fix/per-service-history-duplicate-persistence
Open

Python: fix: prevent superlinear history growth by deduplicating messages in save_messages#7242
PratikWayase wants to merge 1 commit into
microsoft:mainfrom
PratikWayase:fix/per-service-history-duplicate-persistence

Conversation

@PratikWayase

Copy link
Copy Markdown
Contributor

Motivation & Context

  1. Why is this change required?
    The current history providers (InMemoryHistoryProvider, FileHistoryProvider, and RedisHistoryProvider) blindly append messages on every service call without checking if they already exist in the store.
  2. What problem does it solve?
    In looped runs (e.g., AG-UI stateless clients or harness todo loops), the transport passes the full accumulated conversation as input on every request. This caused the entire conversation history to be re-persisted every round, leading to superlinear store growth, corrupted context, repeated tool calls, and massive token bloat (up to 3× the real conversation size).
  3. What scenario does it contribute to?
    Multi-iteration flows where each service call's message list carries the accumulated conversation rather than a delta, particularly when using stateless chat clients over the AG-UI endpoint.
  4. Issue link: Fixes Python: [Bug]: per-service-call history persistence re-appends already-persisted messages #7211

Description & Review Guide

  • What are the major changes?

    • Added a new _get_message_identity helper function that generates a stable identity for a message (using message.id if available, otherwise falling back to a deterministic hash of role + serialized contents).
    • Updated save_messages in InMemoryHistoryProvider, FileHistoryProvider, and RedisHistoryProvider to build a set of existing message identities and filter out duplicates before appending/pushing new messages.
    • Added comprehensive regression tests for all three providers to ensure identical messages are ignored, mixed old/new messages only append the new ones, and messages with the same text but different roles are kept separate.
  • What is the impact of these changes?
    This hardens the system at the lowest storage level. It completely prevents duplicate message persistence regardless of how the middleware or transport layer constructs the message list. This stops token bloat, prevents context corruption, and ensures graceful degradation in long-running loops without requiring changes to the middleware routing logic.

  • What do you want reviewers to focus on?

    1. The stability and correctness of the _get_message_identity fallback logic (ensuring it handles missing IDs and serialization edge cases gracefully).
    2. The RedisHistoryProvider.save_messages implementation, specifically ensuring it correctly fetches existing messages via lrange to build the identity set before pushing only the new_messages via the pipeline.
    3. The new unit tests to ensure they adequately cover the deduplication edge cases.

Related Issue

Fixes #7211

Contribution Checklist

  • The code builds clean without any errors or warnings
  • All unit tests pass, and I have added new tests where possible
  • The PR follows the Contribution Guidelines
  • This PR is linked to an issue and there is no other open PR for this issue (see Related Issue above).
  • This is not a breaking change. If it is a breaking change, add the breaking change label (or add "[BREAKING]" to the title prefix, before or after any language prefix) — a workflow keeps the label and title prefix in sync automatically.

Copilot AI review requested due to automatic review settings July 21, 2026 18:52
@giles17 giles17 added the python Usage: [Issues, PRs], Target: Python label Jul 21, 2026
@github-actions github-actions Bot changed the title fix: prevent superlinear history growth by deduplicating messages in save_messages Python: fix: prevent superlinear history growth by deduplicating messages in save_messages Jul 21, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes a Python history-persistence defect where per-service-call history providers re-persist the full accumulated conversation every round, causing superlinear growth and duplicated context. It introduces message-level deduplication in history providers (in-memory, file, Redis) and adds regression tests to ensure repeated inputs don’t bloat stored history.

Changes:

  • Added a message-identity helper and used it to filter out already-persisted messages before appending/pushing.
  • Updated save_messages behavior across in-memory, file (JSONL), and Redis history providers to skip duplicates.
  • Added unit/regression tests covering identical-message deduplication and mixed old/new message lists.

Reviewed changes

Copilot reviewed 4 out of 4 changed files in this pull request and generated 2 comments.

File Description
python/packages/core/agent_framework/_sessions.py Adds message identity + deduplication to in-memory and file history providers.
python/packages/core/tests/core/test_sessions.py Adds regression tests for deduplication behavior in core providers and looped runs.
python/packages/redis/agent_framework_redis/_history_provider.py Adds message identity + deduplication to Redis history persistence.
python/packages/redis/tests/test_providers.py Adds Redis-provider tests validating deduplication behavior.

Comment on lines +86 to +96
msg_id = getattr(message, "id", None)
if msg_id is not None:
return ("id", msg_id)

try:
# Use to_dict() for a stable, deterministic representation
serialized = json.dumps(message.to_dict(), sort_keys=True, ensure_ascii=False)
return ("content", message.role, serialized)
except Exception:
# Fallback if serialization fails for any reason
return ("content", message.role, str(message.contents))
Comment on lines +1321 to +1325
existing_identities: set[tuple] = set()
if file_path.exists():
with file_path.open("r", encoding="utf-8") as f:
for line in f:
line = line.strip()
Uses the message's ID if available, otherwise falls back to a hash of
its role and serialized contents to prevent duplicate persistence.
"""
msg_id = getattr(message, "id", None)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this checks for id on Message, which does not exist, message_id does.

return
existing = state.get("messages", [])
state["messages"] = [*existing, *messages]
existing_identities = {_get_message_identity(m) for m in existing}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this has the potential to be a major performance hit, recalculating the id every time and especially since id does not exists, it is doing json.dumps every run, for every message!

from redis.credentials import CredentialProvider


def _get_message_identity(message: Message) -> tuple:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this should not be replicated in each history provider implementation, rather a single method should be part of the core and that should be used by all, the problem is likely present in each HistoryProvider.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

python Usage: [Issues, PRs], Target: Python

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Python: [Bug]: per-service-call history persistence re-appends already-persisted messages

4 participants