Information-Theoretic Asymmetric Architecture · v2.0

ASVD-Bridge:
Zero-Overhead Subspace Stabilization

Compressing an 8-Billion Parameter LLM to a 4.67 GB VRAM Footprint at 32+ TPS.

bash
$ pip install qtensor
scroll
4.67 GB
VRAM
Fits Consumer GPUs
32.93 t/s
Throughput
Real-Time Generation
8.03B
Parameters
Base Meta-Llama-3
Step 230
Alpha Release
Actively Training
Zero-Overhead
Inference
Folded ASVD-Bridge
4.67 GB
VRAM
Fits Consumer GPUs
32.93 t/s
Throughput
Real-Time Generation
8.03B
Parameters
Base Meta-Llama-3
Step 230
Alpha Release
Actively Training
Zero-Overhead
Inference
Folded ASVD-Bridge
Architecture Concept

The Pressure Regulator for Neural Networks

Think of an AI model like a massive city water grid. Compressing it shrinks the pipes, which causes the mathematical 'water pressure' (tensor variance) to drop, breaking the downstream valves. The Subspace Bridge acts as a dynamic pressure regulator placed immediately after the compressed pipes to recalibrate the pressure. Because it is a Rank-1 transformation, we permanently weld that regulator directly into the pipe itself after training, costing zero extra compute during inference.

Verified Benchmarks · ASVD-Bridge (Step 230) · NVIDIA RTX 5080
Metric FP16 Baseline ASVD-Bridge Demo
VRAM Footprint ~15.0 GB 4.67 GB ✅
Throughput ~15 t/s 32.93 t/s ✅
Coherence Status Full Pre-trained Early Syntax Healing (Step 230) ⚠️
Hardware Targets

At 4.67 GB, this architecture fits comfortably on base-tier Mac M-Series chips (8GB Unified), standard RTX 3060/4060 GPUs, and lightweight backend server containers.

Live Training Roadmap

Current Checkpoint: Step 230. Semantic convergence expected at Step 5,000. Follow the repository for the final weight release.

Developer Experience

Compress any LLM in
three lines.

QTensor's asymmetric compression pipeline automatically maps entropy to rank, applies AWQ activation scaling to every MLP projection, and injects a learnable Rank-1 Subspace Bridge — all via a single qtensor.compress() call.

compress.py
import qtensor
from transformers import AutoModelForCausalLM
# Load and compress via Information-Theoretic rank allocation
model = AutoModelForCausalLM.from_pretrained("NousResearch/Meta-Llama-3-8B")
model = qtensor.compress(
    model,
    strategy="asymmetric_v2",
    use_triton=True
)
Technical Architecture

Three-Tier Compression Topology

Each compression layer targets a distinct information-theoretic regime, minimizing the KL divergence from the teacher manifold at every precision boundary.

LAYER 01 · ATTENTION

Entropy-Mapped SVD

Shannon entropy over the singular value distribution maps each attention projection to a dynamic rank budget r ∈ [8, 32]. Low-entropy layers get r=8; information-dense layers expand to r=32.

Block-SVD SpLoRA QKV + O
LAYER 02 · MLP

AWQ INT4 MLPs

Activation-Aware Weight Quantization scales salient channels before INT4 packing: W_packed = INT4(W / S) with X_scaled = X × S in the forward pass. Zero catastrophic weight collapse.

INT4 AWQ Gate/Up/Down Triton GEMV
LAYER 03 · BRIDGE

Rank-1 Subspace Bridge

A learnable scalar injected after each attention layer aligns the student's residual stream with the teacher's hidden-state manifold during QAD distillation. Cures precision-boundary variance mismatch.

Rank-1 QAD Loss KL + MSE
🔒 AIR-GAPPED · ZERO-CLOUD · SOVEREIGN AI

Deploy on proprietary
hardware. Your rules.

Custom sub-5GB 8-Billion parameter architectures built for zero-cloud, air-gapped environments. Purpose-built deployments for defense, legal M&A, robotics, and enterprise edge inference.

CTO / ML EngineeringDefense & GovernmentLegal & M&A AIEdge RoboticsPrivate Cloud
Book Architecture Consultation →