Compressing an 8-Billion Parameter LLM to a 4.67 GB VRAM Footprint at 32+ TPS.
Think of an AI model like a massive city water grid. Compressing it shrinks the pipes, which causes the mathematical 'water pressure' (tensor variance) to drop, breaking the downstream valves. The Subspace Bridge acts as a dynamic pressure regulator placed immediately after the compressed pipes to recalibrate the pressure. Because it is a Rank-1 transformation, we permanently weld that regulator directly into the pipe itself after training, costing zero extra compute during inference.
| Metric | FP16 Baseline | ASVD-Bridge Demo |
|---|---|---|
| VRAM Footprint | ~15.0 GB | 4.67 GB ✅ |
| Throughput | ~15 t/s | 32.93 t/s ✅ |
| Coherence Status | Full Pre-trained | Early Syntax Healing (Step 230) ⚠️ |
At 4.67 GB, this architecture fits comfortably on base-tier Mac M-Series chips (8GB Unified), standard RTX 3060/4060 GPUs, and lightweight backend server containers.
Current Checkpoint: Step 230. Semantic convergence expected at Step 5,000. Follow the repository for the final weight release.
QTensor's asymmetric compression pipeline automatically maps entropy to rank,
applies AWQ activation scaling to every MLP projection, and injects a learnable
Rank-1 Subspace Bridge — all via a single qtensor.compress() call.
Each compression layer targets a distinct information-theoretic regime, minimizing the KL divergence from the teacher manifold at every precision boundary.
Shannon entropy over the singular value distribution maps each attention projection to a dynamic rank budget r ∈ [8, 32]. Low-entropy layers get r=8; information-dense layers expand to r=32.
Activation-Aware Weight Quantization scales salient channels before INT4 packing: W_packed = INT4(W / S) with X_scaled = X × S in the forward pass. Zero catastrophic weight collapse.
A learnable scalar injected after each attention layer aligns the student's residual stream with the teacher's hidden-state manifold during QAD distillation. Cures precision-boundary variance mismatch.
Custom sub-5GB 8-Billion parameter architectures built for zero-cloud, air-gapped environments. Purpose-built deployments for defense, legal M&A, robotics, and enterprise edge inference.