LATX: optimize AVX scalar physics patterns - #438
Open
luzeng87 wants to merge 3 commits into
Open
Conversation
Track VEX.128 destinations whose architectural YMM high halves are known to be zero, and materialize those clears only when a 256-bit operation can observe them or before leaving the TB. This removes repeated LASX clear instructions while preserving signal, JIT, TU, and AOT-visible state. Add a standalone JIT, cold-AOT, and hot-AOT semantic test for the deferred state. Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary / 变更说明
This PR depends on #437. Until #437 is merged, GitHub also shows its prerequisite commit in this PR; the GB603-specific changes are the final two commits.
VCOMIS*/VUCOMIS*plusJCCsequences while preserving ordered and unordered floating-point behavior.This change does not modify the AOT cache identity or enable LAT
--optimize-O2.Performance / 性能
3A6000-25G, Geekbench 7
geekbench_avx2, workload 603, GCC 8.3.0-O2 -g, LAT--optimize-O1:The clean branch excludes jump-cache, profiler, AES/VAES, and AOT build-identity changes.
Validation / 验证
meson test -C build64-pr --suite lat-pr-fast: 24 passed, 0 failed.avx128-shuffle.S,avx-sum3-pattern.S,repeat-add.S, andvcomis-jcc.S: native x86, LAT JIT, cold AOT, and hot AOT all passed with non-empty AOT files.Checklist / 检查项
CONTRIBUTING.md. / 我已阅读CONTRIBUTING.md。git commit -s). / 每个提交都包含 DCO 签署。