Found by deploying the RDS log-API path to the PostgreSQL monitoring host and reading what it actually said.
What the operator sees
collection_log:
2026-08-26 09:16:03 | pg_plan_capture | SUCCESS | rows=0 | no new auto_explain plans in the RDS log window
What actually happened
[WARN] RDS plan log unavailable for <target>: User: …assumed-role/<monitor-role>/<instance> is not
authorized to perform: rds:DescribeDBLogFiles on resource: <cluster instance> because no identity-based
policy allows the rds:DescribeDBLogFiles action — plan capture is skipped for this target this cycle
Nothing was read. The log was never opened. And the row says, in as many words, that it was opened and held no plans.
The mechanism
RdsPlanIngestor.IngestAsync catches the failure, warns, and returns 0:
catch (Exception ex) when (ex is not OperationCanceledException)
{
_logger?.LogWarning("RDS plan log unavailable for {Server}: {Message} …", storageName, ex.Message);
return 0;
}
DarlingCollectorRunner.IngestRdsPlansAsync then turns 0 into a positive claim:
return new CollectorRunResult(rows, 0, elapsedMs,
rows == 0 ? "no new auto_explain plans in the RDS log window" : null);
Three states collapse into one sentence, and only the first of them is true:
- The log was read and had no plans — a true negative.
- The AWS call was denied. Nothing was read.
- The cluster was failing over, or the endpoint was refused as a reader.
Tolerating the failure is right — one target's IAM gap must not take the cycle down. Reporting it as a successful empty read is not, and it is a REGRESSION against the route it replaced: the pg_read_file path answers the same situation with PERMISSIONS and a message naming the grant. The managed path, which is the one the fleet is on, is the one that goes quiet.
The warning in the app log is not a substitute. collection_log is where collection health is read, and it currently says this collector is fine.
Fix
Distinguish "could not read" from "read nothing". IngestAsync should surface the failure rather than encode it as a row count — the vocabulary already exists and the sibling route already uses it: SQLSTATE-style classification into PERMISSIONS for AccessDenied, ERROR otherwise, with the AWS message carried through. A cycle that could not look must not claim it looked.
Second defect, same area: IsAwsRds is dead on PostgreSQL targets
The dispatch is:
["pg_plan_capture"] = (r, s, ct) =>
s.Target.IsAurora || s.Target.IsAwsRds
? r.IngestRdsPlansAsync(s, ct)
: r.RunAsync(PgPlanCaptureCollector.Instance, s, ct),
DarlingServerConnector.ConnectPostgresAsync sets Engine, PostgresMajorVersion, PostgresVersionNum, IsAurora and IsInRecovery. It never sets IsAwsRds — that field is only assigned on the SQL Server path, from a T-SQL detection query.
So IsAwsRds is always false for every PostgreSQL target, the second half of that condition is unreachable, and plain RDS PostgreSQL — managed, no filesystem, not Aurora — falls to the pg_read_file route, which cannot work there in principle. It then fails with a message telling the operator to check a grant that would not have helped.
Aurora is unaffected because IsAurora carries it (verified on the fleet: detection reports Aurora: True and the target routes to the RDS API correctly). The gap is exactly the non-Aurora RDS PostgreSQL case, which has no test because there is none in the fleet.
Fix: derive IsAwsRds for PostgreSQL from the endpoint, which RdsEndpoint already parses, rather than leaving a gate that reads as covered and is not.
Found by deploying the RDS log-API path to the PostgreSQL monitoring host and reading what it actually said.
What the operator sees
What actually happened
Nothing was read. The log was never opened. And the row says, in as many words, that it was opened and held no plans.
The mechanism
RdsPlanIngestor.IngestAsynccatches the failure, warns, and returns 0:DarlingCollectorRunner.IngestRdsPlansAsyncthen turns 0 into a positive claim:Three states collapse into one sentence, and only the first of them is true:
Tolerating the failure is right — one target's IAM gap must not take the cycle down. Reporting it as a successful empty read is not, and it is a REGRESSION against the route it replaced: the
pg_read_filepath answers the same situation withPERMISSIONSand a message naming the grant. The managed path, which is the one the fleet is on, is the one that goes quiet.The warning in the app log is not a substitute.
collection_logis where collection health is read, and it currently says this collector is fine.Fix
Distinguish "could not read" from "read nothing".
IngestAsyncshould surface the failure rather than encode it as a row count — the vocabulary already exists and the sibling route already uses it: SQLSTATE-style classification intoPERMISSIONSfor AccessDenied,ERRORotherwise, with the AWS message carried through. A cycle that could not look must not claim it looked.Second defect, same area:
IsAwsRdsis dead on PostgreSQL targetsThe dispatch is:
DarlingServerConnector.ConnectPostgresAsyncsetsEngine,PostgresMajorVersion,PostgresVersionNum,IsAuroraandIsInRecovery. It never setsIsAwsRds— that field is only assigned on the SQL Server path, from a T-SQL detection query.So
IsAwsRdsis always false for every PostgreSQL target, the second half of that condition is unreachable, and plain RDS PostgreSQL — managed, no filesystem, not Aurora — falls to thepg_read_fileroute, which cannot work there in principle. It then fails with a message telling the operator to check a grant that would not have helped.Aurora is unaffected because
IsAuroracarries it (verified on the fleet: detection reportsAurora: Trueand the target routes to the RDS API correctly). The gap is exactly the non-Aurora RDS PostgreSQL case, which has no test because there is none in the fleet.Fix: derive
IsAwsRdsfor PostgreSQL from the endpoint, whichRdsEndpointalready parses, rather than leaving a gate that reads as covered and is not.