CUDA
Posts tagged “CUDA”.
- Five 4K60 Streams Were Fine. The Sixth Was Not.
A paced NVDEC-CUDA-NVENC pipeline found the live-video capacity boundary and showed why a NumPy round trip cost more than the CUDA work.
- The Radio Won't Wait for Your FFT
A paced I/Q replay found where an eight-core SciPy pipeline began missing deadlines and dropping blocks, and how far an RTX GPU moved that boundary.
- I Wrote the Same GPU Operation Six Ways
Six implementations of one row-scoring operation show why a 324× resident GPU kernel becomes 10.4× once host-device transfers are included.
- The Projects Where Data Movement Cost More Than Compute
Examples from game engines, search, distributed processing, deep learning, and ETL where changing data layout or movement mattered more than changing the arithmetic.
- The GPU Collision Detector That Didn't Get Faster
A CUDA collision detector was no faster at the tested scene size because of memory layout, transfer costs, divergent work, and coarse timing.