OpenDLSS-NR in 2026: The Open-Source Project That Broke DLSS 5's RTX-50-Only Lock — What It Means for RTX 40 and AMD Home Labs
TL;DR: A solo developer reimplemented NVIDIA’s DLSS 5 Neural Rendering network in open-source Vulkan — bit-exact — and it runs on RTX 40-series cards NVIDIA said couldn’t have it. The catch: without Blackwell’s NVFP4 hardware, the FP8 fallback costs 39–50% of your frame rate. If you bought a used RTX 4090 for local AI, this is one more reason not to regret it.
| OpenDLSS-NR on RTX 40 | Official DLSS 5 on RTX 50 | AMD RDNA 4 (mod) | |
|---|---|---|---|
| Best for | RTX 40 owners who want the tech today | Buyers who want full-speed neural rendering | Curiosity only |
| Cost of entry | Free (MIT) + the card you own | RTX 5090: $3,822–$5,000 street | RX 9070 XT: $600–$800 |
| The catch | 39–50% frame-rate hit at FP8 | Paying 2026 Blackwell street prices | ~30 fps at 1080p, anti-cheat blocks it |
Honest take: Don’t buy a new GPU because of neural rendering, in either direction. OpenDLSS-NR plus NVIDIA’s own confirmed RTX 40 rollout means the RTX-50-exclusive window is closing on its own — a used RTX 4090 at $2,150–$2,350 remains the better local-AI buy than a $3,800+ RTX 5090 unless you need 32GB of VRAM.
What actually happened
On September 21, 2026, a developer going by maanHimself published OpenDLSS-NR on GitHub: an MIT-licensed Vulkan reimplementation of the neural rendering network inside DLSS 5. The repo hit the Hacker News front page in early October and sits at 809 stars with 65 forks as of October 4, 2026.
The claim that made people look twice is bit-exactness. This is not an approximation or a “similar-looking” filter. The project reconstructs the full 71-block shifted-window transformer plus ViT U-Net that NVIDIA ships, and its intermediate outputs match the original at every processing boundary, byte for byte. VGTimes and ByteIota both independently reported the bit-for-bit parity claim, and the repo documents the verification method: feed the same low-dynamic-range frame, Gaussian noise, and conditioning scalars through both pipelines and diff the buffers.
Two implementation paths ship in the repo:
- Vulkan — the fast path. Uses tensor cores through
VK_KHR_cooperative_matrixand NVIDIA’s cooperative-matrix and FP8 extensions, with FP8 E4M3 activations and fused matrix operations. Requires an Ada Lovelace (RTX 40) or newer NVIDIA GPU on Windows. - WebGPU — a browser port that produces the same bit-exact output without tensor cores at all. Slower, but it runs anywhere a modern browser runs.
What the repo does not contain: NVIDIA’s model weights. The network is 141 MiB of weights that you must supply yourself as a manifest-based directory — the project includes no NVIDIA software and no instructions for extracting anything. It also doesn’t touch DLSS Super Resolution or Frame Generation; this is only the Neural Rendering network, the new piece in DLSS 5.
What DLSS 5 Neural Rendering is, in one paragraph
DLSS 5 launched September 3, 2026 alongside NBA 2K27, and unlike every previous DLSS it is not an upscaler. Neural Rendering takes a finished frame at native resolution and generatively re-renders it: lighting, skin subsurface scattering, material response, and shadow detail get replaced by the output of a learned model running inside the frame loop. NVIDIA markets it as the “GPT moment for graphics” and locked it to RTX 50-series at launch, citing deep optimization for NVFP4 — the 4-bit floating-point format only Blackwell tensor cores execute natively. On a RTX 5090, NVIDIA claims NBA 2K27 runs up to 370 fps at 4K Ultra with ray tracing and the full DLSS 5 stack.
That NVFP4 justification is the part OpenDLSS-NR just stress-tested. The same network, quantized to FP8 — which RTX 40-series runs natively — produces identical output. The exclusivity was never about whether Ada could run the network. It was about how fast.
The real numbers on RTX 40-series
Here is what the network alone costs on an RTX 4070 SUPER, from the project’s own benchmarks. This is inference time for the neural rendering pass only, not total frame time:
| Resolution | Network inference time | Budget left for the game at 60 fps (16.7 ms) |
|---|---|---|
| 768×768 | 2.8 ms | plenty |
| 1920×1080 | 7.8 ms | ~8.9 ms — tight |
| 2560×1440 | 12.6 ms | ~4.1 ms — not realistic |
| 3840×2160 | 29.3 ms | over budget before the game renders a single pixel |
In-game results from the parallel modding scene (which got the official DLL running on Ada by other means) tell the same story from the other end. GameGPU’s analysis puts the FP8 penalty at 39–50% of frame rate on RTX 40 cards. Digital Citizen’s test of an RTX 4090 in Control measured ~130 fps dropping to ~83 fps with neural rendering on — a 36% loss. Borncity’s figures from the German modding community: 57–60 fps down to ~37 fps, a 43% cut.
The problem and the fix, because there is one: the naive pipeline runs Neural Rendering at output resolution after upscaling, which at 4K means paying that 29 ms toll on the full-size frame. A mod called Neural Upstream reorders the passes — neural rendering happens on the internal 1080p frame first, then DLSS Super Resolution upscales the result. TweakTown’s test on an RTX 4080 in Assassin’s Creed Shadows: 31 fps with the normal pipeline, 50 fps with Neural Upstream. A 61% improvement from changing the order of operations, at some cost to fine detail that the upscaler has to reconstruct.
And the kicker: NVIDIA blinked. On September 3 — launch day, with the modding community already running leaked DLLs on Ada — NVIDIA confirmed to Club386 and Notebookcheck that official DLSS 5 support is coming to RTX 40-series “once RTX 50 tuning wraps,” expected after this fall’s model updates. No date. But the exclusivity is now officially temporary.
Will your card run OpenDLSS-NR?
The Vulkan path needs four extensions. Check what your driver exposes before cloning anything:
$ vulkaninfo | grep -iE "cooperative_matrix|shader_float8"
VK_KHR_cooperative_matrix : extension revision 2
VK_NV_cooperative_matrix2 : extension revision 1
VK_EXT_shader_float8 : extension revision 1
If VK_NV_cooperative_matrix2 is missing, you’re on a pre-Ada NVIDIA card or a non-NVIDIA GPU, and the Vulkan path is out — two of the four required extensions are NVIDIA-vendor extensions. RTX 30-series also lacks FP8 tensor-core support (that’s why NVIDIA’s own RTX 40 rollout stops at Ada and excludes Ampere). The build itself is Windows-only for now: Visual Studio 2022+, Python 3, and Node.js for the WebGPU demo.
The weights question is the honest asterisk over the whole project. OpenDLSS-NR is clean-room on code but useless without the 141 MiB of NVIDIA weights, which ship inside NVIDIA’s driver package for RTX 50 owners. The repo deliberately doesn’t tell you how to get them. Whether extracting them from a driver you legally installed constitutes fair use is an open question the project sidesteps entirely — worth knowing before you build a workflow on top of it.
The AMD angle: it runs, and it hurts
OpenDLSS-NR’s Vulkan path won’t run on AMD (NVIDIA vendor extensions, see above), leaving RDNA 4 owners the slow WebGPU route. But a separate project got there another way: developer danielblnc’s “DLSS-NR for AMD” tool wraps NVIDIA’s own nvngx_dlssnr.dll in a compatibility layer so RDNA 4 cards can run the official network in DirectX 12 games.
The results explain why AMD isn’t advertising this. PC Gamer’s test of an RX 9070 XT in Cyberpunk 2077: roughly 30–33 fps at 1080p with neural rendering enabled — on a card that pushes well over 100 fps in that game normally. The developer has improved throughput 26% since the first release, and anti-cheat software blocks the DLL outright, so it’s single-player only. RDNA 4’s WMMA units simply don’t have an FP8 path that competes with tensor cores here; the 9070 XT’s 640 GB/s of bandwidth isn’t the bottleneck, the matrix throughput is. TweakTown separately reports AMD is preparing its own neural rendering answer for RDNA 5 — which, if you own an AMD card today, is the actual thing to wait for. Our ROCm local AI guide covers what RDNA 4 is good at in a home lab.
Does any of this matter for local AI workloads?
Mostly no, with one real exception and one signal worth reading.
The honest “no”: OpenDLSS-NR is not a tool you’ll plug into your inference stack this month. The network is trained on game render output — LDR frames with specific conditioning inputs — and no ComfyUI node or post-processing integration exists as of October 2026. Nothing architecturally stops someone from feeding it arbitrary images through the WebGPU port, but there’s no evidence yet it does anything useful to a Flux render, and without legally distributable weights, nobody can ship that integration cleanly. If you want faster image generation on RTX hardware today, NVFP4 in ComfyUI on RTX 50-series is the proven path.
The exception: if you run a dual-purpose machine — local AI weekdays, gaming weekends — the DLSS 5 situation now directly affects which card that should be, and that’s the next section.
The signal: this is the second time in 2026 that “Blackwell-exclusive” turned out to mean “unoptimized elsewhere, not impossible elsewhere.” NVFP4 model files genuinely need Blackwell (Qwen3.6-27B at NVFP4 doubles throughput on a 5090 precisely because of that hardware), but the DLSS 5 network itself drops to FP8 with zero quality loss and a measurable, bounded speed cost. A 71-block transformer running in 2.8–29.3 ms inside a frame loop on a mid-range Ada card is also a nice existence proof of how much inference headroom consumer GPUs have for small models. The moat is real, but it’s a performance moat, not a capability moat — price your upgrade accordingly. Our sister site aifoss.dev tracks the open-source side of stories like this one.
Does it change the RTX 40 vs RTX 50 buying call?
At the margin, yes — in favor of used Ada.
The case for paying RTX 50 street prices ($3,822–$5,000 for a 5090, per our September tracking) was always: 32GB VRAM, 1,792 GB/s bandwidth, NVFP4, and the exclusive feature stack. The feature-stack leg just got shorter. Neural rendering runs on Ada today via OpenDLSS-NR and mods, officially “this fall-ish” per NVIDIA, at a 36–43% frame cost that Neural Upstream-style reordering claws back. Meanwhile a used RTX 4090 holds at $2,150–$2,350 with 24GB and 1,008 GB/s — still the card we recommend in used RTX 4090 vs new RTX 5090 for everyone who doesn’t specifically need 32GB residency or NVFP4 model files.
What this does not do is make a 40-series card faster. If your workload is NVFP4-quantized models or you want neural rendering at full frame rate, Blackwell is still the only place that exists. The point is narrower: “RTX 50 or you miss DLSS 5” is no longer true, so don’t let it set your budget.
What to actually buy
Prices as of October 2026, all verified in the comparison above or our running price index:
| Your situation | The machine | Price | Where |
|---|---|---|---|
| Own an RTX 40 card, want DLSS 5 tech now | Keep it — OpenDLSS-NR / mods today, official support coming | $0 | GitHub |
| Buying for local AI + gaming, 24GB is enough | Used RTX 4090 | $2,150–$2,350 | Check price |
| Need 32GB, NVFP4, and full-speed neural rendering | RTX 5090 | $3,822–$5,000 | Check price |
| Mid-range dual-use box, 12GB AI ceiling accepted | Used RTX 4070 SUPER | $538–$595 | Check price |
| Undecided — want to test a 5090 workload first | Rented GPU, from $0.25/hr | pay per hour | Vast.ai |
The RX 9070 XT row is deliberately absent: at ~30 fps with neural rendering it’s not a DLSS 5 purchase, and as an AI card its case is unchanged — see RX 9070 XT vs RTX 5060 Ti.
FAQ
Is OpenDLSS-NR legal to use? The code is MIT-licensed, clean-room, and contains no NVIDIA software — that part is uncontroversial. The weights are NVIDIA’s, ship with NVIDIA’s driver, and the project won’t tell you how to obtain them. Running extracted weights locally for personal use sits in the same gray zone as most interoperability modding; redistributing them does not. Nothing here is settled law.
Will NVIDIA kill it? NVIDIA’s response so far has been the opposite: confirming official RTX 40 support rather than lawyering up. A takedown of MIT-licensed clean-room code with no bundled weights would be a hard case, but the repo depending on driver-extracted weights keeps NVIDIA in control of the practical supply chain either way.
Does this bring DLSS 5 to RTX 30-series or GTX cards? No. Ampere lacks FP8 tensor-core support and the required Vulkan extensions, so the fast path stops at Ada. The WebGPU port technically runs anywhere, but “runs” and “runs at playable speed” are different claims — the project doesn’t publish WebGPU frame timings, and tensor-core-free inference of a 71-block transformer per frame will not hit 60 fps on old hardware.
Can I use it to upscale my Stable Diffusion or Flux outputs? Not today. It’s not an upscaler (input and output are the same resolution), it’s trained on game renders, and no integration with ComfyUI or any image pipeline exists as of October 2026. Watch the repo — the WebGPU port makes experimentation easy for anyone with weights.
Should I wait for RTX 50 SUPER instead of buying anything now? Per our standing analysis, waiting for new NVIDIA consumer silicon in this market has not been rewarded — supply constraints keep pushing refreshes right while used Ada prices stay rational. Buy for the VRAM you need now.
Recommended Gear
- Used RTX 4090 24GB — $2,150–$2,350; the dual-use pick this story reinforces
- RTX 5090 32GB — $3,822–$5,000; only if you need NVFP4 and 32GB residency
- RTX 4070 SUPER — $538–$595 used; the card the project’s own benchmarks were run on
- RTX 4080 16GB — the Neural Upstream test card
- RX 9070 XT 16GB — $600–$800; fine AI value, not a DLSS 5 buy
Sources
- Solo Developer Reimplements DLSS 5’s Neural Network in Open-Source OpenDLSS-NR — XenoSpectrum
- OpenDLSS-NR repository (maanHimself) — GitHub
- NVIDIA DLSS 5 Gets Open-Source Vulkan Reimplementation, Runs on RTX 40 GPUs — Wccftech
- Open-source DLSS 5 implementation is bit-for-bit identical to NVIDIA’s — VGTimes
- OpenDLSS-NR: DLSS 5 Neural Network Reversed, Byte for Byte — ByteIota
- NVIDIA locks DLSS 5 to September 3 and RTX 50, launching with NBA 2K27 — Pasquale Pillitteri
- NVIDIA Announces DLSS 5 With 3D-Guided Neural Rendering — MySmartPrice
- DLSS 5 adds realistic shadows and lighting even without hardware RT (FP8 penalty analysis) — GameGPU
- Modder Gets NVIDIA DLSS 5 Neural Rendering Running on RTX 40 Series GPUs (RTX 4090 Control test) — Digital Citizen
- Modders boost DLSS 5 performance by over 50% (Neural Upstream, RTX 4080) — TweakTown
- Nvidia confirms DLSS 5 support for GeForce RTX 40 series cards, but won’t say when — Club386
- Nvidia confirms DLSS 5 is coming to RTX 40-series GPUs, no release date — Notebookcheck
- My RX 9070 XT: Please stop, it hurts us (DLSS 5 on AMD test) — PC Gamer
- DLSS 5 Manager adds RDNA 4 support with new AMD mode — PCGuide
- AMD reportedly prepping its own DLSS 5 answer for RDNA 5 — TweakTown
- RTX 4070 SUPER used market pricing — ResalePrices
Last updated October 4, 2026. Prices and specs change; verify current rates before purchasing.
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →Get the numbers before you buy
New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.