ProBackend
ai hardware custom chips
5 hours ago5 min read

Redesigning AI Server Hardware Architecture for the LLM Era

MatX's $500 million Series B, led by Jane Street and Situational Awareness, funds the MatX One — a purpose-built LLM processor based on a splittable systolic array and an SRAM-HBM hybrid memory design. A look at the funding, the architecture, the competitive field, and the execution risks ahead of a 2027 launch.

Redesigning AI Server Hardware Architecture for the LLM Era

When two former Google hardware leads decided to build a chip startup from scratch, they weren't looking for incremental performance gains. Reiner Pope and Mike Gunter walked away from TPU development in 2023 with a clear, uncompromising thesis: general-purpose accelerators built for traditional deep learning are fundamentally misaligned with the economic and computational realities of large language models. Last month, that foundational bet secured a massive $500 million Series B funding round, thrusting the newly minted competitor into the heavyweight tier of AI infrastructure startups.

The financing, led by Jane Street and Situational Awareness—the investment fund launched by former OpenAI researcher Leopold Aschenbrenner—brings MatX's total capital raised to over $700 million. Despite this influx of cash, the company remains pre-revenue and is racing toward a 2027 product launch. This is a high-stakes, high-risk maneuver that pits a deeply specialized engineering team against Nvidia's entrenched software ecosystem and the brutal physics of advanced semiconductor manufacturing.

The MatX One: An LLM Chip Designed From First Principles

In the announcement accompanying the round, Pope gave the company's processor a name—the MatX One—and sharpened its goal: deliver much higher throughput than any other chip while also achieving the lowest latency, with the explicit ambition of being 10 times better than Nvidia's GPUs at training LLMs and serving them.

What distinguishes the MatX One is not merely aggressive targets but the declared willingness to sacrifice breadth to hit them. Pope wrote that the chip is designed "from first principles with a deep understanding of what LLMs need and how they will evolve," and that the team is willing to give up small-model performance, low-volume workloads, and even ease of programming to build it. That is a genuinely different posture from most AI hardware architecture programs, which chase general coverage across training, inference, classical ML, and HPC workloads in a single part. MatX is betting that the LLM workload class is now large enough, and economically important enough, to deserve silicon that does one thing exceptionally well.

The Systolic Array: A Different Approach to AI Server Hardware Architecture

MatX is designing its chips specifically around the two distinct phases of LLM computation: prefill (processing input tokens) and decode (generating output tokens). The company's central architectural bet is its "splittable systolic array," which attempts to dynamically allocate computational resources between these phases depending on workload demands. According to the company, the design retains the energy and area efficiency that large systolic arrays are known for while still achieving high utilization on the smaller, flexibly shaped matrices that dominate real inference traffic, an area where conventional systolic arrays often sit idle.

The memory subsystem makes a similar attempt at having things both ways. The chip combines high-bandwidth memory (HBM) with a large bank of on-chip SRAM, pairing the low latency of SRAM-first designs with the long-context support that HBM capacity enables. The stated result is higher throughput on LLMs than any announced system while simultaneously matching the latency characteristics of SRAM-first designs. A "fresh take on numerics", the company's arithmetic precision choices, completes the stack. The architecture reflects a broader shift in AI server hardware architecture: optimizing entire systems around model behavior rather than treating accelerator cores as the sole performance lever.

Notably, the company also decided not to build its own network fabric or formulators for the chip, judging those full products to be non-core to its mission. It will instead use third-party formulators and integrate its accelerator directly with other components in the server pipeline. That is a pragmatic de-risking decision for a 100-person startup, and a tacit acknowledgment that in modern AI edge infrastructure and datacenter deployments alike, the accelerator is only one element of a system economics problem.

Capital, Competition, and a Familiar Fundraising Arc

The Series B reads as a signal about how venture money now orbits foundation-model infrastructure. Beyond the Jane Street and Situational Awareness co-leads, participants include Spark Capital, NFDG, Daniel Gross and Nat Friedman's fund, Triatomic Capital, Harpoon Ventures, Stripe co-founders Patrick and John Collison, and individual backers such as Andrej Karpathy and Dwarkesh Patel. MatX also recruited investors across the supply chain, including Alchip and Marvell Technology, partners whose involvement hints at packaging and silicon-integration relationships that matter as much as capital in getting a leading-edge chip to volume production.

TechCrunch previously reported MatX's 2024 funding round valued the company at more than $300 million; the company declined to disclose its latest valuation. The comparison point is hard to ignore: Etched, MatX's closest competitor in the transformers-only chip category, raised a $500 million round at a $5 billion valuation three years ago, also before shipping a chip. Funding is now abundant for this category of AI hardware custom-chip venture; proof is not. Pope previously led AI software development for Google's TPUs before co-founding MatX in 2023, and his résumé gives the effort more architectural credibility than the typical chip startup pitch, but fundraising is not a demonstration of commercial performance.

Execution Risk: Where Specialized Architectures Go to Die or Prove Themselves

MatX's challenge is not simply to produce a fast accelerator. The company says it will finish development and quickly scale manufacturing, targeting tapeout in under a year and initial manufacturing in 2027, with production capacity already reserved. Every step of that plan carries risk: tapeout slippage on advanced nodes, yield ramps, software maturity, and the gravitational pull of Nvidia's CUDA ecosystem, which remains the default reason enterprises buy GPUs.

The deliberate narrowness of the MatX One adds a second axis of risk. Giving up ease of programming and small-model workloads is a defensible engineering trade only if LLM serving patterns remain stable through 2027 and beyond; a significant shift in model architectures could leave a single-purpose chip exposed in a way a general accelerator would not be. That is the broader test for AI server hardware architecture: whether specialized designs can translate promising silicon ideas into reliable, scalable infrastructure that customers adopt not because it is elegant, but because it is dependable and cheaper per token.

For now, MatX occupies an interesting position in the AI hardware architecture intern-to-architect talent market as well as the investor market, a pre-revenue company with a TPU-credentialed team, supply-chain investors, reserved manufacturing, and a 100-person team betting everything on one very specific answer to one very expensive question.

Related coverage: Meta's Silicon Pivot: Creating Custom MTIA Chips to Curb GPU Reliance and AI chip startup landscape.

Sources: TechCrunch (February 24, 2026) and MatX's "MatX One and our Series B" research post.

redesigning ai server hardware architecture for the llm

More blogs