Fix vector/matrix product (@) silently overflowing for narrow value dtypes - #44
Merged
Merged
Conversation
…types _major_matvec/_minor_matvec/_major_matmat/_minor_matmat allocated their output accumulator as values.dtype, the compact per-element storage dtype (e.g. uint16 for single-cell counts), instead of a dtype sized for the accumulated total. Contributions landing in the same output slot wrapped modulo 2**bits instead of promoting, unlike sum() which already accumulates in float64. Promote values and the other operand to their common numpy dtype (np.result_type) before running the kernels, matching the dtype promotion an equivalent dense-array product would get. Fixes #43. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
vector @ VCSRArray/VCSRArray @ vector(and the 2-D@ matrixvariants) allocated their output accumulator usingvalues.dtype— the compact stored-element dtype (e.g.uint16for single-cell counts) — instead of a dtype wide enough for the accumulated total, so results silently wrapped modulo2**bitswhilesum()(which already accumulates infloat64) stayed correct.valuesand the other operand to their commonnp.result_typebefore running the_ops.pykernels, matching the dtype promotion an equivalent dense-array product would get.Test plan
test_matmul_does_not_overflow_narrow_value_dtypeintests/test_ops.py, covering both@directions and both vector/matrix operands, for bothVCSRArrayandVCSCArray, reproducing the issue's uint16 repro.uv run pytest -q— 1662 passed, 109 skipped.uv run ruff check— clean.Fixes #43.
🤖 Generated with Claude Code