Repository navigation
CLR memory corruption in .NET 10 and 11 #134928
Description
Activity
- addeduntriagedNew issue has not been triaged by the area ownerNew issue has not been triaged by the area owner
on Sep 30, 2026 dotnet-policy-service commented
on Sep 30, 2026 ContributorMore actionsTagging subscribers to this area: @agocke
See info in area-owners.md if you want to be subscribed.AI triage:
It's a stale per-thread LoaderAllocator handle for collectible thread statics:
- Each thread keeps its collectible thread-static data alive via a handle in
Thread::pLoaderHandles[tlsIndex], allocated from the type's LoaderAllocator. - On ALC unload,
FreeTLSIndicesForLoaderAllocatorfrees the TLS index, but other threads' handles for it are left in place. FindClearedIndexreuses the index for a thread-static type in a new ALC.- When such a thread exits (e.g. a retiring thread-pool worker),
FreeLoaderAllocatorHandlesForTLSDataresolves the index to the new type and callsFreeHandleon the new ALC with the old ALC's handle. That nulls an unrelated slot in the new ALC's handle table and returns it to the free list, or writes past the table's end.
Objects in that slot get collected while still in use, hence the
0x80131506/ AVs / bogus NREs.A Checked build of main asserts on the repro within ~3 min:
Assert failure: i < GetNumComponents() LoaderAllocator::SetHandleValue LoaderAllocator::FreeHandle FreeLoaderAllocatorHandlesForTLSData Thread::OnThreadTerminateA small deterministic repro (thread X touches a
[ThreadStatic]in collectible ALC1 and stays alive, ALC1 unloads, ALC2 reuses the index, thread X exits) silently nulls the main thread's[ThreadStatic]in ALC2 on 11 RC1.Index reuse came with #99183, so .NET 9+ is affected. Possible fix: store the LoaderAllocator creation number alongside each per-thread handle and skip
FreeHandleon mismatch. Locally that fixes the deterministic repro, and the original repro ran 14 min without crashing.- Each thread keeps its collectible thread-static data alive via a handle in
- removeduntriagedNew issue has not been triaged by the area ownerNew issue has not been triaged by the area owner
on Sep 30, 2026
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsNo status
Description
I ran into a flaky-tests issue on an internal project, where tests were occasionally crashing the CLR for no apparent reason. After trying and failing to debug it, I asked Claude to hammer on it until it could track the problem down, and went to bed. It took 3 hours and an inordinate amount of tokens, but it produced a repro case and blamed the problem on #131267. This appears to be a mistaken identification of the problem, as I'm on SDK 10.0.401, which should contain a fix for this issue.
I've verified that the repro case can cause both 10.0.401 and 11 RC1 to crash with CLR memory corruption issues.
Reproduction Steps
corruption.zip
Build and run this project on .NET 10 or 11. (The problem may exist in earlier versions as well; I haven't checked.) The issue is highly non-deterministic and appears to be related to garbage collection of collectible assemblies, so it may take a while to show up, but it does error out.
Expected behavior
Safe code should never cause CLR corruption.
Actual behavior
This repro case will cause crashes, either with an
ExecutionEngineExceptionor a spuriousNullReferenceException.Regression?
Unknown
Known Workarounds
The issue appears to be related to garbage collection of collectible assemblies. If the ALC is non-collectible, the error does not reproduce. This is not ideal, for obvious reasons.
Configuration
.NET versions 10.0.401 and 11 RC1
Windows 10, x64
Other information
This code is ugly. I'm sorry. It's AI generated. It used to be twice as ugly, but I've cleaned it up some. It's still ugly.