San Marcos, Texas · montoyaraul34@gmail.com
Making large models run on hardware you can actually own.
My core research is SAAQ (Spiking Adaptive Activity Quantization) — compressing large-scale models so they run on consumer GPUs by turning them into spiking neural network representations. I also work on model interpretability, quantization benchmarking, and a modular neuromorphic stack under Limen Neural.
| Project | Description | Stack |
|---|---|---|
| corinth-canal | SAAQ reference loop: telemetry → spiking layer → projector/router → latent calibration | Rust · CUDA |
| xai-dissect | Static structural analysis of Grok-family open-weight MoE checkpoints | Rust |
| grok-ozempic | Ternary SNN-inspired quantization for Grok-scale MoE routing fidelity | Rust |
| magere-brug | SAAQ experiment lab: manifests, recipes, artifact registry | Rust |
| Surrogate_Viz.jl | Symbolic regression + validation dashboards for SAAQ telemetry | Julia |
| XAIDissect_Viz.jl | Interactive Grok-1 MoE atmosphere from xai-dissect reports | Julia · Makie |
| myelin-accelerator | Blackwell-first CUDA kernels for neuromorphic inference | Rust · CUDA |
| agoge-forger | Local-first GPU training forge (QLoRA/LoRA, RTX 5080-aware) | Python · PyTorch |
| combine-for-AI | Neutral benchmark harness for model quantization experiments | Python |
| Dioscuri-Cloud | Cloud ML lab: smoke tests, cost ledger, multi-provider runbooks | Terraform · Ops |
| gaming-telemetry | High-frequency GPU/CPU telemetry → Parquet for SNN training | Rust · NVML |
| spike-viz | Visualize SNN encodings from axon-encoder exports | Python · PyTorch |
| NeuralForge-Memory | Personal RAG + vector DB + MCP memory (Hermes tutor) | Rust · MCP |
Limen Neural is my neuromorphic library org — hard repo boundaries, modular crates you can reuse, dual MIT/Apache-2.0 where applicable.
- Rust infra:
neuromod·axon-encoder·cortex-tensor·nir-rs·myelin-accelerator - Julia research:
TemporalFocus.jl·LiquidCortex.jl·NeuroPulse.jl - Hardware path:
silicon-hdl·silicon-bridge(Q8.8 → FPGA)
I ship research with a multi-agent engineering stack across Grok Build, Codex, Claude (incl. Claude Code), Cursor, Devin, Kilo, OpenCode, Cline, and related agent CLIs (e.g. Mimo, Jules). Agents share durable context via Ogham and Chroma, babysit PRs (CI + review threads), and work in isolated worktrees so parallel edits stay safe. Humans merge — agents prepare, never auto-merge.
Tooling: worktree-hive for issue → PR orchestration with isolated subagents.
- SAAQ — Spiking Adaptive Activity Quantization: frontier MoE models on consumer GPUs
- SNN compression — Low-bit, activity-aware quantization of large transformer MoE architectures
- Symbolic regression — Compact equation discovery from high-dimensional GPU/neuromorphic telemetry
- Explainable MoE — Dissect and visualize expert routing in open-weight models (no weight redistribution)
Actively learning and building with:
Languages: Julia · Rust · Python · CUDA/C++ · HCL
Frameworks: PyTorch · CUDA.jl · SymbolicRegression.jl · ort (ONNX)
Infra: Terraform · Multi-cloud · NVIDIA Blackwell (RTX 5080)
Agents: Grok Build · Codex · Claude · Cursor · Devin · Kilo · OpenCode · Cline
Memory: Ogham · Chroma
- 🔥 SAAQ reference path via
corinth-canal+ Grok-scale quantization / dissect viz - 🧬 Growing the Limen Neural stack: encode → dynamics → NIR → FPGA
- 🤖 Multi-agent PR and worktree workflows across the agent fleet above
- ⚡ Local Blackwell / RTX 5080 kernels and training loops
- 🎓 B.S. in AI Engineering @ WGU — deepening Julia, Rust, Python, CUDA, and HCL in real repos
README updated with Grok Build: Grok 4.5



