Memory Systems
Co-Designed from Silicon to Tokens

Netpreme builds the Memory Processing Unit (MPU) — a new class of hardware and software that gives every GPU in the rack orders of magnitude more memory.

Learn about Technology and Usecases

X-Mem™: Memory Tier Built for AI

Hardware
  1. GPU
    G1HBM10 TB/s · 0.3 TB
  2. G1.5 & G2.ExpX-Mem™3.6 TB/s · 6 TB
Higher Bandwidth
Higher Capacity
Software
GPUX-Mem™GPU KernelsCUDATritonJAXFlashInferX-Mem™ LibraryInference / Training Runtime
  • GPU-Native Memory Tiering Library
  • Transparent to GPU kernels and inference / training runtimes
  • Verified solution for KV/prefix caching, long-context inference, multi-model serving, and optimizer/activation offloading
Benchmarks
  • 7×

    faster TTFT than DRAM offload on prefix-heavy workloads. Benchmark ↗︎

  • 4×

    higher concurrency with the same number of GPUs.

  • 40%

    cheaper inference for long-running background agents over software optimizations. Inference ↗︎

We are hiring

Join Our Team

We’re hiring across Silicon, Systems, and ML with offices in Cambridge, MA and Santa Clara, CA.