Files
opencv-MIRROR/modules/dnn
Pratham Kumar 8602d9d964 Merge pull request #29831 from pratham-mcw:nchw_nchw_opt
dnn: fix eltwise layer memory traversal pattern causing Windows-ARM64 slowdown - #29831

## Problem
- `EltwiseInvoker::operator()` (backs the `Eltwise` layer's SUM/PROD/DIV ops) ran slower on Windows-ARM64 than x64.

## Root cause
- The loop tiled the output into fixed-size blocks and looped over the **channel axis inside each block**.
- In NCHW layout, consecutive channels of the same sample sit far apart in memory (`planeSize * 4` bytes apart).
- So each block jumped between ~256 memory locations per tensor — roughly **768 concurrent strided streams** across the three tensors involved.
- Hardware prefetchers can only track a small, fixed number of streams and Windows-ARM64's prefetcher falls over on this pattern.

## Fix
- Walk the output buffer contiguously and track the current (sample, channel) plane from the offset and clamp each block so it never crosses into the next channel.
- This removes the inner channel loop entirely: one channel finishes before the next starts, so every tensor is read/written sequentially instead of jumping around.
- No behavior change, only the traversal order is different.

## Performance Benchmarks:
<img width="1479" height="208" alt="image" src="https://github.com/user-attachments/assets/0f878757-d988-4564-b166-96d80834c77e" />
2026-09-01 14:26:36 +03:00
..