Skip to content

RDS plan capture reports SUCCESS "no new plans" when the AWS call was DENIED #2633

Description

@erikdarlingdata

Found by deploying the RDS log-API path to the PostgreSQL monitoring host and reading what it actually said.

What the operator sees

collection_log:
2026-08-26 09:16:03 | pg_plan_capture | SUCCESS | rows=0 | no new auto_explain plans in the RDS log window

What actually happened

[WARN] RDS plan log unavailable for <target>: User: …assumed-role/<monitor-role>/<instance> is not
authorized to perform: rds:DescribeDBLogFiles on resource: <cluster instance> because no identity-based
policy allows the rds:DescribeDBLogFiles action — plan capture is skipped for this target this cycle

Nothing was read. The log was never opened. And the row says, in as many words, that it was opened and held no plans.

The mechanism

RdsPlanIngestor.IngestAsync catches the failure, warns, and returns 0:

catch (Exception ex) when (ex is not OperationCanceledException)
{
    _logger?.LogWarning("RDS plan log unavailable for {Server}: {Message} …", storageName, ex.Message);
    return 0;
}

DarlingCollectorRunner.IngestRdsPlansAsync then turns 0 into a positive claim:

return new CollectorRunResult(rows, 0, elapsedMs,
    rows == 0 ? "no new auto_explain plans in the RDS log window" : null);

Three states collapse into one sentence, and only the first of them is true:

  1. The log was read and had no plans — a true negative.
  2. The AWS call was denied. Nothing was read.
  3. The cluster was failing over, or the endpoint was refused as a reader.

Tolerating the failure is right — one target's IAM gap must not take the cycle down. Reporting it as a successful empty read is not, and it is a REGRESSION against the route it replaced: the pg_read_file path answers the same situation with PERMISSIONS and a message naming the grant. The managed path, which is the one the fleet is on, is the one that goes quiet.

The warning in the app log is not a substitute. collection_log is where collection health is read, and it currently says this collector is fine.

Fix

Distinguish "could not read" from "read nothing". IngestAsync should surface the failure rather than encode it as a row count — the vocabulary already exists and the sibling route already uses it: SQLSTATE-style classification into PERMISSIONS for AccessDenied, ERROR otherwise, with the AWS message carried through. A cycle that could not look must not claim it looked.

Second defect, same area: IsAwsRds is dead on PostgreSQL targets

The dispatch is:

["pg_plan_capture"] = (r, s, ct) =>
    s.Target.IsAurora || s.Target.IsAwsRds
        ? r.IngestRdsPlansAsync(s, ct)
        : r.RunAsync(PgPlanCaptureCollector.Instance, s, ct),

DarlingServerConnector.ConnectPostgresAsync sets Engine, PostgresMajorVersion, PostgresVersionNum, IsAurora and IsInRecovery. It never sets IsAwsRds — that field is only assigned on the SQL Server path, from a T-SQL detection query.

So IsAwsRds is always false for every PostgreSQL target, the second half of that condition is unreachable, and plain RDS PostgreSQL — managed, no filesystem, not Aurora — falls to the pg_read_file route, which cannot work there in principle. It then fails with a message telling the operator to check a grant that would not have helped.

Aurora is unaffected because IsAurora carries it (verified on the fleet: detection reports Aurora: True and the target routes to the RDS API correctly). The gap is exactly the non-Aurora RDS PostgreSQL case, which has no test because there is none in the fleet.

Fix: derive IsAwsRds for PostgreSQL from the endpoint, which RdsEndpoint already parses, rather than leaving a gate that reads as covered and is not.

Activity

  1. erikdarlingdata commented on Aug 26, 2026

    @erikdarlingdata
    OwnerAuthor

    Fixed in #2634. An authorization refusal is now PERMISSIONS with a message saying the grant is on the monitoring HOST's IAM role rather than the database login, and that nothing was read; every other failure stays loud. IsAwsRds is derived from the endpoint on the PostgreSQL path, so the dispatch's second half is reachable and plain RDS PostgreSQL routes correctly.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions