Vectorize min + max and add fused minmax - #22759
Conversation
Summary: Replace the iterator-returning `std::min_element` / `std::max_element` implementations of `torch::executor::vec_minf` and `vec_maxf` with four independent value-reduction lanes. A shared compile-time implementation computes only the requested extrema, with a narrowly scoped Clang vectorization hint for NEON/SSE2 targets and no fast-math requirement. Add `torch::executor::vec_minmaxf(const float* x, size_t size, float* min_out, float* max_out)` to compute both extrema together. Wire the per-tensor and both serial/parallel per-token `choose_qparams` paths to the fused helper, eliminating their separate minimum and maximum scans without changing scale/zero-point calculations. Differential Revision: D119391738
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/22759
Note: Links to docs will display an error until the docs builds have been completed. ❌ 5 New Failures, 1 Cancelled Job, 2 Unrelated FailuresAs of commit 7012c19 with merge base f03eeef ( NEW FAILURES - The following jobs have failed:
CANCELLED JOB - The following job was cancelled. Please retry:
BROKEN TRUNK - The following jobs failed but were present on the merge base:👉 Rebase onto the `viable/strict` branch to avoid these failures
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
|
@JakeStevens has exported this pull request. If you are a Meta employee, you can view the originating Diff in D119391738. |
This PR needs a
|
Summary:
Replace the iterator-returning
std::min_element/std::max_elementimplementations oftorch::executor::vec_minfandvec_maxfwith four independent value-reduction lanes. A shared compile-time implementation computes only the requested extrema, with a narrowly scoped Clang vectorization hint for NEON/SSE2 targets and no fast-math requirement.Add
torch::executor::vec_minmaxf(const float* x, size_t size, float* min_out, float* max_out)to compute both extrema together. Wire the per-tensor and both serial/parallel per-tokenchoose_qparamspaths to the fused helper, eliminating their separate minimum and maximum scans without changing scale/zero-point calculations.Differential Revision: D119391738