Agentic AI
Multi-step agentworkloads
Agents accumulate working state across multi-turn conversations and tool calls. That state stays warm between calls without being evicted to storage. X-Mem™ holds the working set at memory-tier latency.
X-Mem™ Technology — Co-Designed from Silicon to Token
The world's first memory processing unit (MPU) chip that enables AI infrastructure with composable memory-to-compute ratios.
AI workloads generates a lot of warm data that needs to be kept resident at memory speed. A new tier of memory is needed.
HBM is too small to hold an agent’s working set. Storage is too slow to serve it. And the working set is not one session’s problem: it grows with every concurrent agent and user, and it outlives the sessions that created it. The working set is a shared resource on the rack.
Netpreme builds the Warm Memory Tier™ that keeps it resident at memory speed.
We’re hiring across Silicon, Systems, and ML with offices in Santa Clara, CA and Cambridge, MA.