Repository navigation
nvidia-drm: don't accumulate override EDID modes in detect (OOM with drm.edid_firmware) - #1430
Open
gianmarcotoso wants to merge 1 commit into
Open
gianmarcotoso wants to merge 1 commit into
gianmarcotoso wants to merge 1 commit into
Conversation
On kernels without drm_connector::override_edid, __nv_drm_detect_encoder() calls drm_edid_override_connector_update() to fetch the override/firmware EDID. That helper also adds the EDID modes to connector->probed_modes. When detect reports the connector as disconnected, drm_helper_probe_single_connector_modes() skips get_modes() and drm_connector_list_update(), so probed_modes is never drained. On the next detect, add_alternate_cea_modes() walks the whole probed_modes list and adds an alternate-clock copy of every CEA mode, including the copies left over from previous calls, so the list doubles on each probe. With drm.edid_firmware set for a DisplayPort connector and the monitor asleep (HPD low), a compositor re-probing on hotplug events grew kmalloc-128 to ~30 GB in about a minute and the system went OOM. Drop any stale probed modes before calling the helper. The modes reported to userspace are built by nv_drm_connector_get_modes() from NVKMS, which already receives the override EDID, so nothing is lost.
gianmarcotoso
force-pushed
the
nvidia-drm-override-edid-probed-modes
branch
from
October 8, 2026 12:40
619c9ec to
b76ab54
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
With an override/firmware EDID set on a connector (
drm.edid_firmware=or debugfsedid_override), every detect on that connector while it is disconnected leaves the override EDID's modes inconnector->probed_modes. Each later probe doubles the number of CEA modes in that list. A compositor re-probing on hotplug events (for example, a DisplayPort monitor going to sleep) can use up all memory within about a minute and hang the system.This PR drops stale probed modes before nvidia-drm calls
drm_edid_override_connector_update()from detect.Root cause
On kernels without
drm_connector::override_edid(the!NV_DRM_CONNECTOR_HAS_OVERRIDE_EDIDpath),__nv_drm_detect_encoder()callsdrm_edid_override_connector_update()to get the override EDID and pass it to NVKMS. The kernel designed that helper for.get_modes(), and it also adds the EDID's modes toconnector->probed_modes.Normally
drm_helper_probe_single_connector_modes()drainsprobed_modesthroughdrm_connector_list_update(). But when detect reportsconnector_status_disconnected, the helper jumps toexitbeforeget_modes()and beforedrm_connector_list_update(), so the modes stay in the list.The next detect calls the helper again.
drm_edid_connector_add_modes()→add_alternate_cea_modes()walks the wholeprobed_modeslist and adds an alternate-clock copy (60 / 59.94 Hz) of every CEA mode it finds, including the leftovers. The CEA mode count therefore roughly doubles on every probe. With a common EDID that has a CTA extension (VICs 3, 4, 16, 63, 97 in my case), about 25 probes are enough to fill 30 GB ofkmalloc-128(sizeof(struct drm_display_mode)).detectis also called once per possible encoder, so a single probe can call the helper more than once.Fix
Before calling
drm_edid_override_connector_update()in__nv_drm_detect_encoder(), free whatever is inconnector->probed_modes. This is safe:get_modes()anddrm_connector_list_update(), so at that pointprobed_modesonly holds leftovers from earlier detect calls.nv_drm_connector_get_modes()fromnvKms->getDisplayMode(). NVKMS already gets the override EDID throughpDetectParams->overrideEdid, so dropping these duplicates changes nothing visible to userspace.NV_DRM_CONNECTOR_HAS_OVERRIDE_EDIDpath never calls the helper and is unchanged.How it showed up
drm.edid_firmware=DP-1:edid/g93sc.bin,DP-2:edid/g93sc.bin,DP-3:edid/g93sc.bin.DPCONN> Zombie? : 1,Lost device), mutter re-probed the connectors. Within ~60 s the system was unresponsive and the kernel log showed:followed by OOM kills of every user process.
Reproduction
The connector does not need a monitor attached. It only needs an override EDID with CTA modes and repeated probes while it is disconnected:
Do not run many more iterations on an unpatched driver: the growth is exponential.
Related
drm_helper_hpd_irq_event(), which runsdetecton every HPD edge. With that change, the growth described here would happen without any userspace probing, so this fix is a prerequisite for it on systems that use an override EDID.Testing
kmalloc-128unchanged (9556 objects before and after).kmalloc-128stays around 1 MB), and the override EDID modes are still reported afterwards.main(615.78.08).