Memory Systems
Co-Designed from Silicon to Tokens
Netpreme builds the Memory Processing Unit (MPU) — a new class of hardware and software that gives every GPU in the rack orders of magnitude more memory.
Learn about Technology and UsecasesAnimated hero concept: a unit of two MPU™ memory nodes and two XPU compute nodes meshed through a scale-up fabric, with data streaming between them; the view then pulls back to show the same unit tiling outward across the rack.
X-Mem™: Memory Tier Built for AI
Hardware
- GPUG1HBM10 TB/s · 0.3 TB
- G1.5 & G2.ExpX-Mem™3.6 TB/s · 6 TB
Higher Bandwidth
Higher Capacity
- G1HBM10–20 TB/s · 300 GB
- G1.5 & G2.ExpX-Mem™ Tier3.6 TB/s · 6 TB
- G2DRAM/PCIe128 GB/s · 1–6 TB
- G3Flash64 GB/s · 20 TB
Higher Bandwidth
Rack-Scale
- G4Storage16 - 64 GB/s · Elastic
Scale-Out
Software
- GPU-Native Memory Tiering Library
- Transparent to GPU kernels and inference / training runtimes
- Verified solution for KV/prefix caching, long-context inference, multi-model serving, and optimizer/activation offloading
Benchmarks
- 7×
faster TTFT than DRAM offload on prefix-heavy workloads. Benchmark ↗︎
- 4×
higher concurrency with the same number of GPUs.
- 40%
cheaper inference for long-running background agents over software optimizations. Inference ↗︎


Join Our Team
We’re hiring across Silicon, Systems, and ML with offices in Cambridge, MA and Santa Clara, CA.