Add production Qwen3.8-27B support - #498
Conversation
🏗️ Architecture Diff
qwen3_5_vl (hybrid-qwen-vl) / embedding — 18 change(s)Op summary: 32 → 47 nodes --- base
+++ head
@@ -1,14 +1,29 @@
Gather
Constant
Equal
+Equal
+Not
+Or
Unsqueeze
+Reshape
+Cast
+Reshape
Cast
Constant
CumSum
Constant
Sub
Constant
+CumSum
+Constant
+Sub
+ReduceSum
+Add
+Where
+Constant
Clip
+Shape
+Reshape
CastLike
Constant
ShapeAdded nodes:
Modified attributes:
Connectivity changes:
qwen3_5_vl (hybrid-qwen-vl) / vision_encoder — 435 change(s)Op summary: 247 → 271 nodes --- base
+++ head
@@ -1,134 +1,233 @@
+Cast
Reshape
Conv
Reshape
+Constant
+Gather
+Constant
+Gather
+Constant
+Gather
+Mul
+Mul
+Constant
+CumSum
+Constant
+Constant
+Pad
+ReduceSum
+Constant
+Constant
+Range
Slice
-Squeeze
-Slice
-Squeeze
-Slice
-Squeeze
-Mul
-Mul
-ReduceMax
-Scan
Shape
-Constant
-Gather
-Constant
-Squeeze
+ConstantOfShape
+Constant
+Mul
+ScatterElements
+Constant
+CumSum
+Gather
+Sub
+Gather
+Gather
+Mul
+Mod
+Constant
+Div
+Mod
+Constant
+Div
+Div
+Mod
+Constant
+Div
+Constant
+Mod
+Constant
+Mul
+Add
+Constant
+Mul
+Add
+Cast
+Cast
+Cast
+Cast
+Mul
+Sub
+Div
+Mul
+Sub
+Div
+Floor
+Floor
+Cast
+Cast
+Constant
+Constant
+Add
+Min
+Constant
+Add
+Min
+Sub
+Unsqueeze
+Sub
+Unsqueeze
+Constant
+Mul
+Constant
+Mul
+Add
+Add
+Add
+Add
+Cast
+Gather
+Gather
+Gather
+Gather
+Sub
+Mul
+Add
+Sub
+Mul
+Add
+Sub
+Mul
+Add
+CastLike
+Add
+Unsqueeze
+Unsqueeze
+Concat
+Gather
+Gather
+Squeeze
+Squeeze
+Gather
+Gather
+Gather
+Gather
+Concat
+Concat
+Constant
+Equal
+Compress
+Unsqueeze
+Concat
+LayerNormalization
+Transpose
+MatMul
+Add
+Split
+Reshape
+Reshape
+CastLike
+CastLike
+Unsqueeze
+Unsqueeze
+Split
+Split
+Mul
+Mul
+Sub
+Mul
+Mul
+Add
+Concat
+Mul
+Mul
+Sub
+Mul
+Mul
+Add
+Concat
+Reshape
+Reshape
+Shape
+Squeeze
+Constant
Constant
Range
Unsqueeze
Unsqueeze
-Less
+GreaterOrEqual
+Cast
+ReduceSum
+Constant
+Sub
+Unsqueeze
+Unsqueeze
+Equal
+CastLike
+CastLike
+Where
+Unsqueeze
+Unsqueeze
+Unsqueeze
+Unsqueeze
+Attention
+Squeeze
+Transpose
+MatMul
+Add
+Add
+LayerNormalization
+Transpose
+MatMul
+Add
+Gelu
+Transpose
+MatMul
+Add
+Add
+Reshape
+LayerNormalization
+Transpose
+MatMul
+Add
+Gelu
+Transpose
+MatMul
+Add
+LayerNormalization
+Transpose
+MatMul
+Add
+Split
+Reshape
+Reshape
+CastLike
+CastLike
+Unsqueeze
+Unsqueeze
+Split
+Split
+Mul
+Mul
+Sub
+Mul
+Mul
+Add
+Concat
+Mul
+Mul
+Sub
+Mul
+Mul
+Add
+Concat
+Reshape
Reshape
Shape
-Slice
-Constant
-Concat
-Reshape
-Compress
-CastLike
-Add
-Slice
-Squeeze
-Slice
-Squeeze
-Slice
-Squeeze
-Mul
-Mul
-ReduceMax
-Scan
-Shape
-Constant
-Gather
-Constant
-Squeeze
+Squeeze
+Constant
Constant
Range
Unsqueeze
Unsqueeze
-Less
-Reshape
-Shape
-Slice
-Constant
-Concat
-Reshape
-Compress
-Gather
-Gather
-Squeeze
-Squeeze
-Gather
-Gather
-Gather
-Gather
-Concat
-Concat
-Slice
-Squeeze
-ReduceMax
-Scan
-Shape
-Constant
-Gather
-Constant
-Squeeze
-Constant
-Range
-Unsqueeze
-Unsqueeze
-Less
-Reshape
-Shape
-Slice
-Constant
-Concat
-Reshape
-Compress
-Constant
-CumSum
-Constant
-Constant
-Pad
-LayerNormalization
-Transpose
-MatMul
-Add
-Split
-Reshape
-Reshape
-CastLike
-CastLike
-Unsqueeze
-Unsqueeze
-Split
-Split
-Mul
-Mul
-Sub
-Mul
-Mul
-Add
-Concat
-Mul
-Mul
-Sub
-Mul
-Mul
-Add
-Concat
-Reshape
-Reshape
-Shape
-Squeeze
-Constant
-Constant
-Range
-Unsqueeze
-Unsqueeze
GreaterOrEqual
Cast
ReduceSum
@@ -159,89 +258,14 @@
MatMul
Add
Add
-Reshape
-LayerNormalization
-Transpose
-MatMul
-Add
-Gelu
-Transpose
-MatMul
-Add
-LayerNormalization
-Transpose
-MatMul
-Add
-Split
-Reshape
-Reshape
-CastLike
-CastLike
-Unsqueeze
-Unsqueeze
-Split
-Split
-Mul
-Mul
-Sub
-Mul
-Mul
-Add
-Concat
-Mul
-Mul
-Sub
-Mul
-Mul
-Add
-Concat
-Reshape
-Reshape
-Shape
-Squeeze
-Constant
-Constant
-Range
-Unsqueeze
-Unsqueeze
-GreaterOrEqual
-Cast
-ReduceSum
-Constant
-Sub
-Unsqueeze
-Unsqueeze
-Equal
-CastLike
-CastLike
-Where
-Unsqueeze
-Unsqueeze
-Unsqueeze
-Unsqueeze
-Attention
-Squeeze
-Transpose
-MatMul
-Add
-Add
-LayerNormalization
-Transpose
-MatMul
-Add
-Gelu
-Transpose
-MatMul
-Add
-Add
-LayerNormalization
-Reshape
-Transpose
-MatMul
-Add
-Gelu
-Transpose
-MatMul
-Add
-Unsqueeze
-Concat
+LayerNormalization
+Reshape
+Transpose
+MatMul
+Add
+Gelu
+Transpose
+MatMul
+Add
+Unsqueeze
+ConcatAdded nodes:
Removed nodes:
Modified attributes:
Connectivity changes:
Initializer changes:
Legend: ⚪ No change · 🔵 Minor (attrs/inits) · 🟡 Moderate (nodes added/removed) · 🔴 Major (interface changed) |
There was a problem hiding this comment.
Pull request overview
Adds production support for pinned Qwen/Qwen3.8-27B by reusing the existing Qwen3.5 hybrid 3-model VL architecture (decoder + embedding + vision encoder), expanding the runtime/metadata integration surface and adding reduced-real + Olive Q4 validation assets.
Changes:
- Extend Qwen VL embedding + vision pipelines to support packed image-then-video feature ordering and keep the processor boundary float32 (cast at the graph boundary).
- Replace Scan-based Qwen3-VL vision position embedding / rotary / cu_seqlens logic with a packed-stream coordinate approach.
- Add reduced-real pinned fixture + goldens + Olive Q4_K_M recipe and wire in additional ORT GenAI auto-export mappings/tests for
qwen3_5*.
Reviewed changes
Copilot reviewed 19 out of 19 changed files in this pull request and generated 2 comments.
Show a summary per file
| File | Description |
|---|---|
| tests/qwen38_real_weight_test.py | Adds unit + integration coverage for the reduced-real fixture and Olive Q4 package validation. |
| tests/integration_test.py | Pins Qwen3.8 config/processor and adds mixed image/video pipeline parity coverage. |
| testdata/golden/vision-language/qwen3_8-27b-reduced.json | New L4 golden for reduced-real logits top-k validation. |
| testdata/golden/vision-language/qwen3_8-27b-reduced_generation.json | New L5 golden for reduced-real cached generation validation. |
| testdata/cases/vision-language/qwen3_8-27b.yaml | Adds a (CI-skipped) case descriptor for Qwen3.8-27B VL. |
| src/mobius/tasks/_vision_language_3model.py | Forces pixel_values input dtype to float32 and casts to model dtype inside the vision graph. |
| src/mobius/models/qwen35.py | Updates MTP-related documentation/comments to reflect separate optional drafter packaging. |
| src/mobius/models/qwen35_test.py | Adds Qwen3.8 alias/contract tests (registry IDs, MTP classification, scatter semantics, float32 boundary). |
| src/mobius/models/qwen_vl.py | Updates Qwen3-VL embedding scatter to handle packed image+video ordering and token IDs. |
| src/mobius/integrations/ort_genai/auto_export.py | Maps qwen3_5* model types to ORT GenAI qwen3_5 and includes qwen3_5_text in VL model-type sets. |
| src/mobius/integrations/ort_genai/auto_export_test.py | Adds coverage for qwen3_5 model-type resolution, packed vision processor emission, and metadata emission without runtime gates. |
| src/mobius/components/_qwen3_vl_vision.py | Reworks Qwen3-VL vision pos-embed interpolation/rotary/cu_seqlens to avoid control-flow subgraphs. |
| src/mobius/_registry.py | Updates registry test model ID for qwen3_5 to Qwen3.8-27B (and retains qwen3_5_vl test IDs). |
| src/mobius/_configs/_vision_defaults.py | Plumbs video + vision boundary token IDs from HF configs into vision defaults. |
| examples/olive/qwen3_8-27b/validate_reduced_checkpoint.py | New pinned reduced-real range-fetch validator for parity, save/load, and Q4 package audit. |
| examples/olive/qwen3_8-27b/requirements.txt | Adds minimal dependencies for the Olive validation recipe. |
| examples/olive/qwen3_8-27b/README.md | Documents the reduced-real + Olive validation workflow and runtime waivers. |
| examples/olive/qwen3_8-27b/optimize.py | Adds decoder-only Q4_K_M quantization workflow and package assembly logic. |
| examples/olive/qwen3_8-27b/inference.py | Adds direct ORT generation helper for reduced Qwen3.8 hybrid-VL packages. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
cfbbc7a to
ae4fd70
Compare
54433fd to
05a478a
Compare
Treat the official checkpoint as a pinned Qwen3.5 hybrid VL alias, preserve image/video processor semantics, classify MTP as a separate optional drafter, and make packed vision coordinates CUDA-safe without Scan subgraphs. Add deterministic real-processor parity, reduced-real FP32/FP16/BF16 validation, cached generation goldens, and an Olive Q4_K_M package recipe with recurrent-gate stability safeguards. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchu@microsoft.com>
Isolate Olive caches so recurrent-gate exclusions are always honored, use a portable CUDA vision graph around the ORT 1.26 PackedMHA defect, and validate the final Q4 package with CUDA reload plus deterministic cached CPU generation. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchu@microsoft.com>
Emit the faithful qwen3_5 GenAI model type without consulting downstream runtime support, retain processor metadata in Q4 packages, validate local processor reload, and remove the static Olive recipe that could not encode required graph-derived exclusions. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchu@microsoft.com>
Document that repeated identical ORT 1.26 CUDA MatMulNBits runs on the exact assembled package can diverge or produce non-finite logits, while CPU cached generation remains the semantic acceptance path. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchu@microsoft.com>
Map the extracted qwen3_5_text subtype back to faithful qwen3_5 multimodal metadata and select the packed Qwen PatchImage processor pipeline for normal config and local exports. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchu@microsoft.com>
Omit the incompatible production tokenizer from the 256-token reduced fixture while retaining and independently reloading the pinned image and video processor metadata. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchu@microsoft.com>
Omit out-of-range production generation IDs from the token-ID-only fixture and require empty variant directories so stale tokenizer assets cannot survive reruns. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchu@microsoft.com>
Merge source provenance and the token-ID-only input contract into the Q4 manifest, and carry processor waivers through package assembly. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchu@microsoft.com>
Load the example's sibling inference module under a Qwen-specific name so prior Nemotron tests cannot poison Python's module cache in combined CI runs. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchu@microsoft.com>
Construct the absent-video mask entirely in the ONNX graph and preserve the original boolean image mask for media feature selection. Add runtime coverage for image-only embedding scatter when video_token_id is unset. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchu@microsoft.com>
Compute packed vision coordinates once, replace quadratic media lookup with boundary scatter and prefix sum, and reuse patch-local indices for frame boundaries. Simplify equivalent bilinear interpolation so Qwen3.5-VL stays below the deterministic graph-node regression threshold while preserving real image/video parity and CUDA semantics. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchu@microsoft.com>
05a478a to
3758328
Compare
Performance Comparison
|
Preserve qwen3_5_text for decoder-only exports while selecting qwen3_5 for multimodal vision and embedding packages, including local-config exports whose composite config was unwrapped. Cover both generated package topologies without adding a runtime capability gate. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchu@microsoft.com>
Drop the standalone reduced-checkpoint Olive recipe and its directly coupled tests and golden fixtures. Core Qwen3.8 model, processor, metadata, and parity coverage remain unchanged. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchu@microsoft.com>
Keep the qwen3_5 registry test surface pinned to Qwen/Qwen3.5-2B. Qwen3.8 remains covered through its independent pinned configuration and case data instead of replacing canonical Qwen3.5 coverage. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchu@microsoft.com>
State that the pinned 55.6 GB checkpoint exceeds hosted CI resources and that no reduced or quantized private fixture is committed. This replaces the obsolete claim that removed reduced-real and Olive assets are run manually. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchu@microsoft.com>
0722ddb
into
justinchuby-add-nemotron-35-lightning
- add production support for pinned `Qwen/Qwen3.8-27B@1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0` as a checkpoint alias of the existing dense Qwen3.5 hybrid vision-language architecture - preserve the exact 64-layer 48-DeltaNet/16-full-attention text schedule, 27-layer vision encoder, image/video processor contracts, mixed-batch scatter order, dtype boundaries, and optional separately packaged MTP drafter - preserve canonical `qwen3_5` registry regression coverage on `Qwen/Qwen3.5-2B`; Qwen3.8 uses its own pinned case/config assertions instead of replacing that mapping - emit topology-faithful ORT GenAI metadata without a runtime capability gate: decoder-only packages keep `qwen3_5_text`, while multimodal Qwen3.5/Qwen3.8 packages emit `qwen3_5` - remove the standalone Qwen3.8 Olive example, its coupled reduced-real test module, and its private reduced golden fixtures; core model, processor, metadata, and parity coverage remain unchanged Stacked on #487 final skill-split head `a8cd77570bca980861eeadab9e6fa077464f0138`. Exact PR head: `59af1a51f1a162e9444a9be1ea5785ca4ec0ef0b`. Tracks #483. - the real processor from the pinned official revision drives deterministic tiny-model HF/ONNX pipeline parity for image-only, video-only, and a mixed two-row batch with opposite placeholder order: image max abs `0.00697723`, video `0.00411959`, mixed `0.00677243`; cosine is above `0.99999` and final argmax matches - exact raw-config assertions cover the 48/16 hybrid schedule, text/vision dimensions, processor token contracts, and intentional separate MTP packaging - packed-vision coordinates are emitted once and shared by interpolation, rotary, and frame-boundary consumers; media ownership uses boundary `ScatterElements` + `CumSum` in O(patches + media) - deterministic Qwen3.5-VL benchmark: `428 -> 467` top-level nodes (`+9.1%`, model size unchanged), below the unchanged 10% blocker threshold - topology-specific metadata regression coverage generates both package shapes: decoder-only `qwen3_5_text` and multimodal `qwen3_5` - post-removal YAML/coverage/focused suite: `987 passed, 227 skipped`; diff lintrunner clean; independent review found no stale reduced/Olive claims or dangling references - canonical non-integration suite before the deletion-only follow-up: `3962 passed, 61 skipped` The pinned official checkpoint is 55.6 GB and exceeds hosted CI storage and GPU memory; no reduced or quantized private fixture is committed. The Qwen3.8 case records this exact `ci_skip_reason` rather than claiming retained executable evidence. Earlier reduced-real and Olive development artifacts were removed and are not part of the final PR validation surface. Current head `59af1a51f1a162e9444a9be1ea5785ca4ec0ef0b`: - Benchmark run `32314325378` is complete and successful (base, head, and comparison). - CI run `32314325553` is still active. Lint, affected-model detection, L1, L3, and lintrunner have completed successfully; the Linux/Windows test matrices are in progress, while Integration (fast), L4, and L5 are queued. - Architecture Diff run `32314325423` is in progress. Historical run `31870831200` belongs to pre-removal head `067cb489599687b74f96574de20133cb96e016a7`, not the final head. It is retained only as historical context: its three red jobs reproduced on the exact #487 base run `31869640946`, but it is not cited as current-head CI evidence. PR #498 remains ready for maintainer review; auto-merge is off. --------- Signed-off-by: Justin Chu <justinchu@microsoft.com> Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Summary
Qwen/Qwen3.8-27B@1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0as a checkpoint alias of the existing dense Qwen3.5 hybrid vision-language architectureqwen3_5registry regression coverage onQwen/Qwen3.5-2B; Qwen3.8 uses its own pinned case/config assertions instead of replacing that mappingqwen3_5_text, while multimodal Qwen3.5/Qwen3.8 packages emitqwen3_5Stacked on #487 final skill-split head
a8cd77570bca980861eeadab9e6fa077464f0138. Exact PR head:59af1a51f1a162e9444a9be1ea5785ca4ec0ef0b. Tracks #483.Validation
0.00697723, video0.00411959, mixed0.00677243; cosine is above0.99999and final argmax matchesScatterElements+CumSumin O(patches + media)428 -> 467top-level nodes (+9.1%, model size unchanged), below the unchanged 10% blocker thresholdqwen3_5_textand multimodalqwen3_5987 passed, 227 skipped; diff lintrunner clean; independent review found no stale reduced/Olive claims or dangling references3962 passed, 61 skippedL4/L5 waiver
The pinned official checkpoint is 55.6 GB and exceeds hosted CI storage and GPU memory; no reduced or quantized private fixture is committed. The Qwen3.8 case records this exact
ci_skip_reasonrather than claiming retained executable evidence. Earlier reduced-real and Olive development artifacts were removed and are not part of the final PR validation surface.CI
Current head
59af1a51f1a162e9444a9be1ea5785ca4ec0ef0b:32314325378is complete and successful (base, head, and comparison).32314325553is still active. Lint, affected-model detection, L1, L3, and lintrunner have completed successfully; the Linux/Windows test matrices are in progress, while Integration (fast), L4, and L5 are queued.32314325423is in progress.Historical run
31870831200belongs to pre-removal head067cb489599687b74f96574de20133cb96e016a7, not the final head. It is retained only as historical context: its three red jobs reproduced on the exact #487 base run31869640946, but it is not cited as current-head CI evidence.PR #498 remains ready for maintainer review; auto-merge is off.