mirror of
https://github.com/opencv/opencv.git
synced 2026-09-11 04:43:22 -05:00
dnn: fix eltwise layer memory traversal pattern causing Windows-ARM64 slowdown - #29831 ## Problem - `EltwiseInvoker::operator()` (backs the `Eltwise` layer's SUM/PROD/DIV ops) ran slower on Windows-ARM64 than x64. ## Root cause - The loop tiled the output into fixed-size blocks and looped over the **channel axis inside each block**. - In NCHW layout, consecutive channels of the same sample sit far apart in memory (`planeSize * 4` bytes apart). - So each block jumped between ~256 memory locations per tensor — roughly **768 concurrent strided streams** across the three tensors involved. - Hardware prefetchers can only track a small, fixed number of streams and Windows-ARM64's prefetcher falls over on this pattern. ## Fix - Walk the output buffer contiguously and track the current (sample, channel) plane from the offset and clamp each block so it never crosses into the next channel. - This removes the inner channel loop entirely: one channel finishes before the next starts, so every tensor is read/written sequentially instead of jumping around. - No behavior change, only the traversal order is different. ## Performance Benchmarks: <img width="1479" height="208" alt="image" src="https://github.com/user-attachments/assets/0f878757-d988-4564-b166-96d80834c77e" />